ingrid.fyi India's software hiring trends, explained.
Sign in to ingrid.fyi with Google
Your Google account
name@gmail.com
Continue with Google
Home / IT services firms / Data engineers in the services world
On this page
Data engineers in the services world A role across the services world, every Company Tier together, Accenture left out. Accenture posts more data engineering than the rest of the services world put together, and its data work is described with the rest of its hiring on its own. The short answer Every analytics dashboard and every AI model a client buys needs its data gathered, cleaned and moved first, and data engineers build the pipelines and platforms that do it. In the services world they are a solid role, above their weight at product companies, because so much of a services firm's data work is moving a client's data onto a modern cloud platform. The giants, like IBM, Infosys and TCS, hold the largest share of the work, though at a little below their services weight. The role is most concentrated at the quiet foreign majors, like CGI and NTT DATA, and among the specialist firms, where firms that do nothing but data and analytics sit, like YASH Technologies, Sigmoid and Fractal, with the build firms, like EPAM and Luxoft, and the consulting arms, like PwC and Deloitte, a little above the rate. The postings require Python above all, with SQL, Spark, and a cloud data platform such as Databricks, Snowflake or Azure Data Factory. Data engineering fell with the rest of services hiring into 2026 but held up better than most roles, and its share recovered in July to September 2026 as the mid-sized firms, like Virtusa and Happiest Minds, and TCS came back. The door is narrow. Senior openings outnumber mid-level ones, and junior openings are well under the services rate. Where the work sits The giants hold the largest share. Company Tier 1, The IT services giants, posts more data engineering than any other tier, but at a little below its services weight. IBM, Infosys and TCS hire the most, and CGI, among the foreign majors, close behind them. The quiet foreign majors and the specialist firms are where it weighs most. In Company Tier 3, The quiet foreign majors, data engineering is well above its services weight, mostly at CGI, NTT DATA and Sopra Steria. In Company Tier 8, Specialist firms built around one skill, it is as high, because the index includes firms whose whole business is data and analytics. YASH Technologies posts the most, and at Sigmoid, CloudSufi, Tredence, Impetus and Fractal data engineering is a large part, or most, of what they post. The build firms and the consulting arms sit a little above the rate. At the build firms the work is almost all EPAM's, building clients' data platforms. At the consulting arms it is PwC and Deloitte, building the data behind the analytics and AI their consultants recommend, and data engineering is the main role at KPMG. Below the rate. The mid-sized firms hire data engineers at a little below their services weight. The small shops, like Wissen Technology and Sahaj Software, the BPO and KPO firms, like EXL and Movate, the infrastructure firms, like Kyndryl and Unisys, and the engineering-services firms, like ACL Digital and L&T Technology Services, hire them at well under it. What the work is

Most data engineering in services is platform work, building the pipelines and data platforms that move and transform a client's data. A small part is centred on the database itself, in SQL engines such as Presto and Trino and in Oracle, and that part is largest at the small shops.

The postings require Python above everything else, with SQL. Spark, usually through PySpark, is required in many of them, and so is a cloud, Azure most often, then AWS and Google Cloud. The platforms named most are Databricks, Azure Data Factory and Snowflake, with Airflow for scheduling pipelines. Many postings also require Java or Scala, and a good share require streaming with Kafka.

The emphasis differs by kind of firm.

The giants lean to Spark at scale, with Scala, Java and Databricks. The quiet foreign majors lean to the Microsoft data stack, Azure Data Factory, PySpark and Databricks. The mid-sized firms and the consulting arms lean to Snowflake and Airflow, and the consulting arms to Redshift and AWS. The build firms still require much of the older Hadoop and Hive stack beside Spark, for the large data platforms they rebuild.

None of the data postings ask for experience with language models. But a sizeable share involve feeding AI and machine-learning systems, and that share was highest in April to June 2026.

Data engineering is spread across cities much as services hiring is, with Bengaluru far ahead, then Hyderabad, Pune and Chennai. Kolkata holds slightly more of it than of services hiring generally.

What is changing

These readings follow the record quarter by quarter, from October to December 2025, the third quarter of the financial year 2025-26, to July to September 2026, the second quarter of 2026-27. A longer record will show whether they last.

Data engineering fell, then recovered a littleIts openings dropped steeply into January to March 2026 and again into April to June, like services hiring as a whole, then rose in July to September 2026 while services hiring kept falling. Its share of services hiring is now well above its share at product companies, where data engineering has eased. The giants left and came back in partIBM and Infosys, the largest data employers in the winter of 2025, posted almost no data jobs after March 2026, but TCS posted many again in July to September 2026. The mid-sized firms and the specialist firms carried 2026YASH Technologies grew its data hiring through 2026, EPAM's peaked in April to June 2026, and Virtusa posted the most of its record in July to September 2026, so in that quarter the mid-sized firms posted more data openings than any other tier. The modern cloud stack gainedSnowflake, Airflow and BigQuery were required in a larger share of data openings in 2026 than in the winter of 2025. Spark held steady.
Persistently open roles

Data jobs are persistently open less often than services jobs overall. When one is posted again, its first and last postings are typically about three weeks apart, and rarely more than two months.

The build firms and the giants keep data jobs open longest. EPAM, NTT DATA, IBM and Infosys post the same data job again far more often than the services average, while the mid-sized firms and the small shops rarely do. Streaming and NoSQL are the hardest skills to find. The skills found most often in persistently open data jobs are Cassandra, Azure Stream Analytics, DynamoDB, Elasticsearch and the Parquet file format, the tools of real-time and very large data.
The door

Data engineering's door is narrow. Junior openings are well under the services rate and entry-level openings almost absent, while senior openings outnumber mid-level ones. Data platforms carry a client's most valuable information, and firms hire people who have built one before.

The consulting arms take juniors most readily, PwC above all, then the mid-sized firms, at Virtusa and ValueMomentum, and the giants, at Infosys and IBM. The build firms, the small shops and the engineering-services firms post no junior data openings. At the build firms and the small shops the role is strongly senior. The specialist firms are the most mid-level of the larger employers, the likeliest door for an engineer with a couple of years of experience.
In one line If you build data pipelines in Python, SQL and Spark on Databricks, Snowflake or Azure, the services world wants you more than product companies do and kept wanting you through 2026, but it wants experience, and the door opens mostly at the mid level and above. Terms used on this page Data platform The system that gathers a company's data from many sources and keeps it ready for reports and AI models. ETL and ELT Extracting data from where it is made, transforming it and loading it into a platform, in either order. Spark and PySpark An engine for processing very large data across many machines, and its Python interface. Databricks, Snowflake and BigQuery Cloud data platforms on which companies store and process their data. Azure Data Factory and Airflow Tools for scheduling and running data pipelines. Hadoop and Hive An older generation of tools for storing and querying very large data. Streaming and Kafka Processing data continuously as it arrives, and the most common system for carrying it. Cassandra, DynamoDB and Elasticsearch Databases built for very large or fast-changing data, outside the usual table-based kind. Presto and Trino Engines for querying data where it already sits, across many stores.
Back to the story

This page reads the openings. The story of the companies is on IT services firms in India.

Privacy Terms Refunds and cancellation Shipping and delivery © 2026 ingrid.fyi · Payments by Razorpay
You're browsing as a guest. Sign in free to follow links for five minutes, once an hour.