ingrid.fyi India's software hiring trends, explained.
Sign in to ingrid.fyi with Google
Your Google account
name@gmail.com
Continue with Google
Home / Product companies / III · Data platforms and AI / Analytics and data providers / Data integration and governance platforms
On this page
Data integration and governance platforms Copying data out of every system into one place, cleaning it, and mapping what a company has You order a phone from an online electronics store. That one order is recorded in several places. The store's app saves it in its order database. The payment system records your payment. The delivery partner's system tracks the parcel. If you send the phone back, the return is logged in the store's customer-support tool.

Now say the store is based in Bengaluru. On Monday morning, its head of sales asks a simple question: how many phones did we sell in Pune last week, and how many came back? The answer needs data from all four systems. Each system is separate, and each stores its data in its own format. A large company runs hundreds of systems like these.

So companies copy the data out of all these systems into one place built for answering questions. That place is called a data warehouse. A data warehouse is a database too. But it is kept apart from the databases the apps run on, so that heavy questions don't slow the apps down. (More on data warehouses in Market Segment 7.1, Databases and data warehouses, in Industry Vertical 7, Database and storage companies.) The answer to the Pune question then appears on a dashboard, a screen of charts built from the warehouse's data. (More on dashboards in Market Segment 8.2, Business intelligence (BI) and analytics platforms.)

One phone order is recorded in four separate systems; pipelines copy, clean and map that data into a warehouse, which feeds the dashboard that answers how many phones sold in Pune and how many came back.One phone order is recorded in four separate systems; pipelines copy, clean and map that data into a warehouse, which feeds the dashboard that answers how many phones sold in Pune and how many came back.

This Market Segment is about the work in between. The data has to be copied into the warehouse, cleaned when it arrives, and mapped, so that people know what the company has. The code that copies data from one system to another is called a pipeline.

The software engineers who build and run pipelines are called data engineers. Together with the analysts who use the data, they make up a company's data team. The work is unglamorous, but it takes up most of a data team's time. That is why it became an industry. The data team is the main buyer of the products in this Market Segment.

Copying data is not the same as connecting applications. Connecting applications means that an action in one app sets off an action in another. For example, when HR adds a new hire, the IT system creates the new hire's email account. Boomi and Workato sell this kind of connection. (More on them in Market Segment 10.2, Integration platforms (iPaaS) and managed file transfer, in Industry Vertical 10, ERP and business automation.)

There are two ways to buy the tools in this Market Segment. Informatica sells all of it in one suite, the oldest of the large integration suites. Many newer companies each sell one piece, and the data team joins the pieces together. The industry calls this set of separate tools the modern data stack. Lately, the pieces have been coming back together. dbt Labs has merged with Fivetran, and larger technology companies have bought Informatica, Reltio and Confluent.

Every company named here has posted software engineering jobs in India. Famous companies that don't actively hire software engineers in India are left out. Most of the companies here have their Indian teams in Bengaluru, and one of them, Hevo Data, was founded there.

This Market Segment has three sub-segments:

Moving it: copying data from where it is created into the warehouse, in batches or as a live stream Knowing what you have: the catalogue, the quality checks, and one agreed record of each customer, product or supplier The industrial data platform: the same work for a refinery or a plant, where much of the data comes from sensors on machines

Keep reading, free

Three more sections are on this page: Moving it, Knowing what you have, and The industrial data platform. Sign in to read them here, in full.

Continue with Google
  • Moving it
  • Knowing what you have
  • The industrial data platform
Moving it ETL and ELT (the two orders of extract, transform and load), change data capture, streaming, connectors, embedded integration, low-code data engineering

Take the online store from the start of this Market Segment. Before anyone can answer the Pune question, the orders, payments, deliveries and returns have to reach the warehouse.

A pipeline does three jobs:

Extractit pulls the data out of the source system, the system where the data was first created. Transformit cleans the data, fixes formats and joins tables together. Loadit puts the data into the warehouse. In the old order, ETL, data was cleaned on its way into the warehouse; in ELT, the order today, raw data is loaded first and transformed inside the warehouse.In the old order, ETL, data was cleaned on its way into the warehouse; in ELT, the order today, raw data is loaded first and transformed inside the warehouse.

The old order was extract, transform, load, or ETL. The data was cleaned on its way in, because space in the warehouse was expensive. Cloud warehouses made storage cheap and computing easy to add. So many teams now load the raw data first and transform it inside the warehouse. This order is called ELT.

Pipelines copy data in one of two ways. A batch pipeline copies everything new at fixed times, for example once a night. A streaming pipeline sends each change the moment it happens.

Some of the products below are no-code. The data team builds a pipeline by clicking and dragging on a screen, without writing code. Low-code products work the same way, but let engineers add code where they need it.

Ready-made pipelines. The first group sells pipelines that are already built, so that the data team doesn't have to write each one by hand.

Fivetranwith engineers in Bengaluru, made the copying a product. It offers hundreds of ready-made connectors. A connector is code that knows how to pull data out of one particular app or database. Fivetran's connectors copy the data into the warehouse and keep it current, without anyone writing the pipeline. dbt Labs merged with Fivetran in June 2026. Its tool, dbt, is widely used to write the transformations inside the warehouse: the T in ELT. Matillionwith engineers in Hyderabad, sells a visual pipeline builder. It loads the data and reshapes it inside the cloud warehouse. Hevo Datafounded in Bengaluru, sells a no-code pipeline much like Fivetran's. Its buyers are data teams in many countries. Prophecyfrom Palo Alto with its engineers in Bengaluru, is the low-code version for data engineers themselves. Its visual editor produces real code underneath, which engineers can read and change. It has moved on to generating the pipeline from a written description.

Selling the connectors themselves

Nexlawith engineers in Bengaluru, and CData Software sell connectors as a product. Their connectors form a layer that makes any source look like any other. Refoldformerly Cobalt, with its engineers in Bengaluru, sells the same idea to a different buyer: software companies, not data teams. A software company's product often has to connect to the other apps used by the businesses that buy it. Refold first sold an embedded integration platform, which a software company builds into its own product to make those connections. Now it sells AI agents that build the connections.

Moving data as it happens. Two companies here stream the data instead of copying it in batches.

Confluentnow part of IBM, with engineers in Bengaluru, is the company behind Apache Kafka. Kafka carries events between a company's systems in real time. An event is a record that something happened, such as an order placed, a click or a payment. Kafka is one of the most widely used systems of its kind. It is free, open-source software that began inside LinkedIn, and its creators started Confluent. Companies built around open-source software usually earn money by running it as a cloud service, and by selling support and extra features. Striimwith engineers in Chennai, reads each change out of a database as it is written and streams it onward. This is called change data capture. With it, the warehouse is only seconds behind the source system, instead of a night behind. A batch pipeline copies everything new at fixed times, so the warehouse can be a night behind; a streaming pipeline sends each change as it happens, so it is only seconds behind.A batch pipeline copies everything new at fixed times, so the warehouse can be a night behind; a streaming pipeline sends each change as it happens, so it is only seconds behind.

One company here works differently. Shakudo, with engineers in Bengaluru, does not move data. It sells what it calls an operating system for data and AI teams. In plain words, it is one platform on which a team installs and runs all the data and AI tools it uses. It runs on the buyer's own systems.

Airbyte and SnapLogic are not covered in this Market Segment. Neither are Astronomer, Dagster and Prefect, which sell orchestration tools. Orchestration tools schedule a company's pipelines and run them in the right order. Talend, now part of Qlik, is not covered either, and neither is Qlik.

Knowing what you have Data catalogues, data quality and observability, data integrity, master data management

Back to the online store from the start of this Market Segment. The data has reached the warehouse. But can anyone find it, and can they trust it? A large company's warehouse can hold thousands of tables. Which of them is the real table of orders? And if the dashboard shows zero sales in Pune, is that true, or did last night's pipeline fail?

This sub-segment is the map and the checks. It has two groups of companies. The newer ones sell the map and the checks. The older ones sell data integrity.

The map and the checks

Alationwith engineers in Chennai, sells a data catalogue. A catalogue is a searchable list of every table a company has, who owns it and what it means. DataHubwith engineers in Bengaluru, is the company behind the open-source catalogue of the same name. Like Kafka, the catalogue was born inside LinkedIn. Acceldatawith its engineering in Bengaluru, sells data observability. It watches pipelines and tables for three problems that can silently break a dashboard:
  • Freshness: did today's data arrive on time?
  • Volume: did the expected number of rows arrive?
  • Schema drift: did a source system change a column's name or type without warning?
Revefiwith engineers in Bengaluru, does the same job with an AI model. The model reads the warehouse and works out why something failed, and why the warehouse bill is as high as it is. Cloud warehouses charge for the computing each query uses, so a badly written query shows up on the bill. Altimate AIsells an AI data engineer: an AI tool that writes the transformations and documents them.

Data integrity: the older half. Data integrity means making sure the records themselves are complete, correct and consistent.

Informaticanow part of Salesforce, sells the whole of this Market Segment in one suite: integration, quality, catalogue and master data. (More on Salesforce in Market Segment 14.1, CRM software, in Industry Vertical 14, CRM and sales tech.) Preciselywith engineers in Delhi, Bengaluru and Pune, sells data quality, enrichment and address data. Enrichment means filling in details a record is missing. Together, these make records trustworthy. Addresses show why this is hard: the same flat in Pune can be written in many different ways. Reltionow part of SAP, with engineers in Bengaluru, sells master data management. It builds one agreed record of each customer, product or supplier from every system that holds a version of it. The industry calls this the golden record. For example, the store's app may list a customer as "Priya S.", the payment system as "PRIYA SHARMA", and the support tool by phone number alone. Master data management works out that these are one person. (More on SAP in Market Segment 10.1, ERP and business accounting software, in Industry Vertical 10, ERP and business automation.) Verdantissells the same kind of product, built for one kind of record: a plant's spare parts and suppliers. (More on Verdantis in Market Segment 30.1, Supply chain planning software, in Industry Vertical 30, Supply chain and logistics.) The store's app, the payment system and the support tool each hold a different version of the same customer; master data management works out they are one person and builds one golden record.The store's app, the payment system and the support tool each hold a different version of the same customer; master data management works out they are one person and builds one golden record. The pipeline that watches itself.

Acceldata, Revefi and Altimate AI show where this sub-segment is heading. AI models now write the transformations, fill in the catalogue and watch the checks. Finding out why a number is wrong is one of a data team's largest costs. The goal is for the platform to report the problem before anyone asks. Back to the online store: the platform would flag the zero for Pune before the head of sales saw it.

Collibra, Atlan, Monte Carlo and Ataccama are not covered in this Market Segment.

You're reading as a guest. Sign in free to follow links for five minutes, once an hour. The industrial data platform Industrial data operations (industrial DataOps), engineering and manufacturing data integration

At the online store, the data came from apps and databases. At an oil platform, a refinery or a factory, much of it comes from machines. Sensors on pumps, turbines and pipes send readings every second. Maintenance records sit in one system and engineering drawings in another. This last sub-segment does the whole job of this Market Segment for data like this.

Cognitewith engineers in Bengaluru, sells a platform that gathers the sensor readings, maintenance records and engineering drawings of an oil platform or a plant. It joins them into one connected picture of the site. Engineers can then ask questions of a refinery, the way the store's head of sales asks questions of the warehouse. Schneider Electric has agreed to buy Cognite. (More on Cognite in Market Segment 37.4, Predictive maintenance and industrial IoT platforms, in Industry Vertical 37, Manufacturing tech.) eQ Technologicwith engineers in Pune, connects the engineering data of aerospace and defence manufacturers. Their designs, parts and test results live in dozens of systems, and eQ Technologic brings them together. (More on the design software itself in Market Segment 37.2, CAD and PLM software, in Industry Vertical 37, Manufacturing tech.) That is the work between the systems that create data and the people who ask questions of it. The data is copied into one place, checked, and matched into one agreed record of each thing. The question may be about phones sold in Pune or about a pump in a refinery, but the work is the same.
Who they hire

Who these companies hire, and for what, is on What analytics and data providers hire for.

Privacy Terms Refunds and cancellation Shipping and delivery © 2026 ingrid.fyi · Payments by Razorpay
You're browsing as a guest. Sign in free to follow links for five minutes, once an hour.