ingrid.fyi India's software hiring trends, explained.
Sign in to ingrid.fyi with Google
Your Google account
name@gmail.com
Continue with Google
Home / GCCs / Data engineers in the GCCs
On this page
Data engineers in the GCCs A role across the GCC collection, all ten Parent Industries together. The short answer Data engineers build the pipelines that move a company's data from where it is made to where it is used. That means the transactions, the sales, the claims and the trades, cleaned, joined and loaded into the warehouses and lakes that reports, risk models and AI are built on. They are the fourth-largest single role in the GCC world, close behind GenAI, and far larger at the GCCs than at product companies, because a bank, a retailer or an insurer runs on its own data in a way most software companies do not. The role is the most evenly spread across industries of any large GCC role. Walmart, JPMorganChase and Societe Generale post the most, but the centres where data is the heart of the work are the asset managers' and investment banks' centres, Macquarie, Fidelity International, Goldman Sachs and Franklin Templeton among them, and Saint-Gobain's systems centre. Through 2026 the role's share of GCC hiring fell, as it did at product companies, and its tools moved from Hadoop to Databricks. It is a mid-career role. Entry-level openings are very rare, junior ones below the GCC rate, and the door opens most at the healthcare companies, the carmakers and the insurers, at centres like Carelon, ADM and MetLife. Where the work sits Every large industry hires it at about its usual weight. Unlike Java, cloud or GenAI, data engineering is not pulled towards one kind of parent. The banks, like JPMorganChase and Societe Generale, hold the most, at a little above their weight, the retailers, like Walmart and Tesco, the next most, also a little above, and the healthcare companies and the carmakers and manufacturers hire it at about their GCC weight. The banks hold the most, and the investment side hires it hardest. Parent Industry 1, Inside the banks and financial institutions, posts more data openings than all the other industries together, led by JPMorganChase, Societe Generale, Wells Fargo, NatWest Group and Morgan Stanley. The centres where data is the largest part of hiring are the asset managers and investment banks, Macquarie, Fidelity International, Goldman Sachs, Franklin Templeton, BNP Paribas and Morgan Stanley, whose business is markets data. (More on it in the banks' hiring, in Parent Industry 1.) The retailers are close behind. Walmart posts more data openings than any other centre, and Tesco, lululemon and Best Buy hire data engineers above their weight. (More on it in the retailers' hiring, in Parent Industry 3.) The healthcare companies hire it at their weight, at Carelon, Optum India and Bristol Myers Squibb. (More on it in the healthcare companies' hiring, in Parent Industry 5.) The carmakers and manufacturers hire it through their systems centres. In Parent Industry 4, Inside the manufacturers and carmakers, Saint-Gobain hires data engineers as one of its main roles, and Ecolab, ADM and Cargill hire them beside their business platforms. The media groups hire it above their weight. In Parent Industry 9, Inside the media, entertainment and publishing groups, Warner Bros. Discovery and Warner Music Group hire data engineers above the GCC rate. Below their weight. The insurers hire data engineers at well under their GCC weight, almost all at MetLife. The telecom operators hire them at about their weight, mostly at TMUS Global Solutions, and the hotels and airlines through Marriott and Delta Air Lines. The logistics operators, like Cardinal Health and Maersk, and the energy majors, like bp and Woodside Energy, hire almost none. By mandate, data is everywhere. The centres Building the parent's apps and platforms, like Walmart, JPMorganChase, Societe Generale and Tesco, hold the most, a little above their weight. The centres that took on The whole engineering function, like Wells Fargo, NatWest Group, Morgan Stanley and Carelon, hold most of the rest, at their weight. Unlike the other large roles, the centres Running the parent's own systems, like Saint-Gobain, Ecolab and Mattel, also hire data engineers at their full weight, because a company's business systems produce the data that must be moved. The data and AI centres, like ADM's and Gap's, are the most data-heavy of all. What the work is

Nearly every data opening is for building data platforms and pipelines. A small minority are database-centred, tuning and running the databases themselves.

Python is required in nearly every posting and SQL in most. A large minority also require Java, and a growing share Scala. Spark appears in about every other posting, with Kafka, Airflow, Databricks, Hadoop and Snowflake behind it, on AWS, Azure and Google Cloud. Most postings involve batch processing and the orchestration of pipelines, and many require loading cloud data warehouses, handling streams of data as they arrive, and building the data behind AI and machine learning. A notable share require data governance and data quality, keeping track of what data exists, who may use it and whether it can be trusted.

The tools follow the industry.

The banks carry an older estate beside the new, with the ETL tools Ab Initio and DataStage and the job schedulers AutoSys and Control-M, alongside Apache Iceberg tables and the Neo4j graph database. The retailers lean to Google Cloud's BigQuery, to Cassandra and Redis for fast lookups, and to the Hadoop tools of a large online store. The healthcare companies lean to AWS, with Redshift, Glue and EMR, and to Microsoft Fabric and Tableau. The carmakers and manufacturers lean to a newer cloud data stack, with Azure Data Factory, dbt, Snowflake and Power BI. The insurers lean to Microsoft's data tools, Azure Synapse and Data Factory, with Informatica. The media groups lean to Flink, for processing streams of data as they arrive.

The whole-function centres carry more of the older estate, with Teradata, DataStage, HBase and AutoSys. The systems centres lean most to the newer stack, with Fivetran, dbt and Snowflake.

Bengaluru and Mumbai hold a larger share of data engineering than of GCC hiring as a whole. Chennai and Pune hold clearly less.

What is changing

These readings follow the record quarter by quarter, from October to December 2025, the third quarter of the financial year 2025-26, to July to September 2026, the second quarter of 2026-27. A longer record will show whether they last.

Data hiring rose and fell with the GCCs, but its share slidData openings doubled from October to December 2025 to January to March 2026, held through April to June, and fell back in July to September 2026 about as sharply as GCC hiring did. Its share of GCC hiring fell through the record, most steeply after October to December 2025, much as data engineering's share fell at product companies. Hadoop gave way to DatabricksHadoop was required in a notable share of data postings in October to December 2025 and in far fewer by July to September 2026. Databricks rose over the same period, and Scala grew in every quarter. The employers shiftedWalmart's data openings fell sharply in July to September 2026 while Tesco's doubled. The insurers' data hiring grew in every quarter, from almost none. Juniors grew in the last quarterThe junior share of data openings was highest in July to September 2026.
Persistently open roles

Data jobs stay open slightly more often than GCC jobs overall. When one is posted again, its first and last postings are typically about three weeks apart, and nine in ten are within two months.

The media groups, the carmakers and manufacturers and the retailers keep data jobs open longest. The healthcare companies rarely post one twice. Commonwealth Bank, Morgan Stanley and Walmart keep a large share of their data jobs open again and again. NatWest Group, Optum India and Carelon rarely do. Staff-level data jobs stay open longest, and the few junior ones fill fastest. The skills found most often in persistently open data jobs are those of the older big-data and streaming estate, Hadoop, Kafka, NiFi, Avro, Cassandra, Teradata, Elasticsearch and Splunk, together with dbt.
The door

Data engineering is a mid-career role. Entry-level openings are very rare and junior openings below the GCC rate. Mid-level openings take a clearly larger share than across the GCCs, and staff-level openings a little less.

The door opens most at the healthcare companies, the carmakers and manufacturers and the insurers, each of which hires juniors into data far more readily than the banks or the retailers. The juniors are at Wells Fargo, Walmart, JPMorganChase, ADM and Carelon. In proportion, Carelon and Wells Fargo take juniors most readily. Tesco, NatWest Group, Morgan Stanley and Commonwealth Bank post no junior data openings. The whole-function centres take juniors more readily than the platform centres, and the platform centres hold most of the staff-level work.
In one line If you build data pipelines in Python and SQL with Spark, every kind of GCC wants you, the asset managers and investment banks most of all, though the door opens mainly once you have a few years behind you. Terms used on this page Mandate The kind of work a parent company gives its engineering centre in India, such as building its apps and platforms, running its own systems, or testing software built elsewhere. Asset manager A firm that invests other people's money, such as pension savings or mutual funds.
Back to the story

This page reads the openings. The story of the companies is on Global capability centres in India (GCCs).

Privacy Terms Refunds and cancellation Shipping and delivery © 2026 ingrid.fyi · Payments by Razorpay
You're browsing as a guest. Sign in free to follow links for five minutes, once an hour.