ingrid.fyi India's software hiring trends, explained.
Sign in to ingrid.fyi with Google
Your Google account
name@gmail.com
Continue with Google
Home / Data engineers / Skills
On this page
The data engineer's skills Job posts for this role want an engineer who builds the pipelines and platforms that move, clean and store a company's data, so that analysts can report on it and models can learn from it. Where a backend engineer builds the application that creates records, a data engineer collects those records from every application, shapes them into one usable form, and keeps them flowing. Nearly all the job posts describe pipeline and platform work. A small share describe a database-centred job, closer to a database developer, where the work is tables and queries on a cloud database without the big data tools. The language and the query

Two things are named more than anything else in this role. SQL is the first, because every stage of a pipeline is a query: pulling records out of a source, joining them, reshaping them, and loading them into a warehouse. Python is the second, and it is the glue that calls the tools, moves the files and writes the transformations that SQL cannot express. Java is third, and Scala fourth, and both mark the heavier end of the work, because Spark itself is written in Scala and the largest pipelines at the marketplaces, streaming services and banks are built in it. Scala and Java are named most at the consumer-data companies and the GCCs, and least at Accenture, where the work is mostly SQL and PySpark. Shell scripting is assumed, and C/C++ appears only in big tech.

Getting the data in: ETL and orchestration

A pipeline starts by reading data from wherever it is made. Application databases, files that partners send, events from an app, logs from servers and records from a vendor's API all arrive in different shapes and at different times. ETL, extract, transform and load, is the name for pulling the data out, cleaning and reshaping it, and writing it where it belongs, and it is the third most-named skill in the role. ELT is the newer order, load first and transform inside the warehouse, and it appears in a smaller share, mostly at the product companies on Snowflake and Databricks. Data Transformation, Data Cleansing and Data Wrangling describe the middle step. Change Data Capture is the technique for picking up only what changed in a source database since the last run.

The tools that do this work follow the cloud. Azure Data Factory is the most-named ETL tool, and it leads in client work and at Microsoft. AWS Glue is Amazon's equivalent. Informatica and Ab Initio are the older enterprise tools, still named at the banks and in services for the systems built with them, and dbt is the newer tool for transformations written in SQL inside the warehouse. ETL and orchestration together are tagged on a large share of job posts.

Orchestration is the scheduling of all these steps so they run in the right order, at the right time, and recover when one fails. Airflow is the tool named most, and it is strongest at the consumer-data companies and the GCCs. Composer is Google's hosted Airflow, and AutoSys is the older enterprise scheduler still named at the banks.

Processing at scale: batch and Spark

Once the data is too large for one machine, it is processed across many, and Spark is the tool for that in nearly every job post that names one. It is the most-named tool after SQL and Python, and at the consumer-data companies it is named in nearly every job post. PySpark is Spark driven from Python, and it is what Accenture and the services firms ask for most, while Scala is how Spark is driven at the marketplaces and the banks. Spark SQL runs queries over Spark. Hadoop is the older platform Spark grew out of, with Hive for querying it, HDFS for storing it and MapReduce as its original way of processing, and all four are still named at the GCCs and the consumer-data companies where the platforms were built on them. EMR is Amazon's hosted Spark and Hadoop. Batch processing is tagged on a large share of job posts, and on nearly all at the consumer-data companies.

Processing as it happens: streaming

Some data cannot wait for the nightly run. A viewer's clicks during a live match, a shopper's search on a marketplace, a payment, a sensor reading from a machine, all have to be processed within seconds. Kafka is the tool that carries these streams, and it is the most-named streaming tool by a wide margin. Flink processes a stream as it flows, Spark Streaming does the same with Spark, and Amazon Kinesis, Pub/Sub and Azure Stream Analytics are the clouds' own streaming services. Streaming is tagged on a fair share of job posts, and the share tells you the employer: it is on most job posts at the consumer-data companies and in big tech, a fair share at the GCCs with the retailers the heaviest, and few at Accenture.

Storing the data: warehouses and lakehouses

The shaped data has to be stored where analysts and models can query it fast. Data Warehousing is named as a skill in its own right, and the warehouse itself is one of the cloud platforms. Databricks is the most-named, and it leads in client work, where moving a client's data onto it is much of what Accenture and the services firms do. Snowflake is second and is strongest at the data companies and the healthcare GCCs. BigQuery on GCP, Redshift on AWS and Azure Synapse and Microsoft Fabric on Azure are the clouds' own warehouses, and Teradata is the older one still named at the banks. Together the cloud warehouse and lakehouse are tagged on a large share of job posts.

A lakehouse keeps the raw files in cheap storage and puts a table layer over them, and the formats for that layer appear by name. Delta Lake is Databricks' format and the most-named, with Iceberg and Hudi as the open alternatives, and Parquet, Avro and ORC as the file formats underneath. Medallion Architecture is the pattern of bronze, silver and gold tables, raw to cleaned to ready, that most lakehouses follow. Presto/Trino queries data across many stores at once, and Druid, ClickHouse and Pinot are the databases built for very fast queries on streams of events, named at the streaming and advertising companies.

Beside the warehouse sit the ordinary databases. SQL Server, PostgreSQL and Oracle are the relational ones named, in small shares, and Oracle marks the banks' older systems. MongoDB, Cassandra, DynamoDB, Redis and Elasticsearch are the stores for records that do not fit tables, and they are tagged as a group on a minority of job posts, most at the consumer-data companies and in big tech.

Keeping the data right: quality and governance

A pipeline that delivers wrong data is worse than one that delivers none, and many job posts ask for the work of keeping it right. Data Quality means checks that the records are complete, consistent and fresh, and Data Governance means knowing where every piece of data came from, who may see it, and how long it is kept. Both are tagged on a fair share of job posts, and most at the GCCs of the banks and healthcare groups, where regulators ask those questions. Data governance tools, catalogues and lineage trackers, are described in the text of job posts rather than named.

Serving the data: reports and models

The pipeline exists so that someone can use the data. At one end are the analysts, and a fair share of job posts want the data engineer to build or support the reports too. Power BI is the reporting tool named most, with Tableau behind it, and this overlap is strongest in big tech and at the consulting arms. At the other end are the machine learning teams, and the ML overlap is tagged on a fair share of job posts, most at the consumer-data companies, the data companies and the GCCs, where the data engineer builds the feature pipelines that feed the models and sometimes the MLOps around them. Pandas and NumPy appear in the job posts closest to that side.

Cloud and containers

Nearly every data platform now lives in the cloud, and cloud is the most common extra ask in the whole role, on most job posts. AWS and Azure are named about equally, and GCP is third, with the order set by the employer: AWS leads at the data companies, the consumer-data companies and the GCCs, Azure leads in client work and at Microsoft, and GCP is strongest at the consumer-data companies. Kubernetes and Docker are named in a minority of job posts, less than in backend roles, because most pipeline tools are hosted services rather than containers the engineer runs. Terraform and CloudFormation appear where the engineer also describes the platform as code, most at the consumer-data companies.

Build pipelines

A data pipeline is code, and it changes as sources and reports change, so it is tested and released through a build-and-release pipeline like any other code. Azure DevOps is the pipeline named most, followed by GitHub Actions, Jenkins and GitLab CI. This is tagged on a fair share of job posts, most at the data companies and the GCCs, and least at Accenture, where the release process is usually the client's.

The application around the pipeline

A fair share of job posts also want some application development, in Java and Spring, .NET, React or Node.js, and the share is highest at the data companies, in big tech and at the GCCs. It means the data engineer also builds the service or screen that exposes the data, or sits on a team where the pipeline is part of a product. In the ordinary data engineering job it is a plus.

The development process

Underneath all of it sits the ordinary craft of building software in a team: Git for source control, pull requests and reviews, issue tracking, and a rhythm of scheduled runs that are monitored and fixed. Job posts count these as given and rarely list them as skills.

Reading the mix

A data engineer who writes strong SQL and Python, can build an ETL pipeline and schedule it in Airflow or Azure Data Factory, processes large data in Spark or PySpark, has worked on Databricks or Snowflake, understands Kafka well enough to handle a stream, and can run it all on AWS or Azure, meets the core of nearly every job post. The variations belong to the employer. The data companies want the platform itself built as a product, with Kubernetes, GitHub Actions and the most application development. The consumer-data companies want the heaviest engineering, Spark in Scala, Kafka, Flink, Hadoop, Delta Lake and streaming on nearly every job post. The GCCs of banks, retailers and healthcare groups want Spark pipelines with Airflow, Kafka at the retailers, and the most Data Quality and governance. Big tech is Microsoft's own stack, Azure, Microsoft Fabric, Azure Synapse and Power BI. Accenture and the services firms, the widest door by far, want SQL, PySpark and ETL to move clients' data onto Databricks and Snowflake, at the mid level. Across all of them, the engineer who is strong in SQL first and Spark second is the one every job post describes.

Who hires data engineers in India
Privacy Terms Refunds and cancellation Shipping and delivery © 2026 ingrid.fyi · Payments by Razorpay
You're browsing as a guest. Sign in free to follow links for five minutes, once an hour.