Data Pipeline Tools: 12 Compared, and How to Choose (2026)

Compare 12 data pipeline tools by the job they do, with selection criteria, where each fits, pricing models and a first pipeline you can run locally.

Data pipeline tools are the software that moves data from where it is produced to where it is used, and changes it on the way. No single product does the whole job well. The market splits into five kinds of tool, and most teams end up running two or three of them together:

  • Managed ingestion (ELT): copies data out of SaaS apps and databases into a warehouse. Fivetran, Airbyte.
  • Event streaming: a durable log that many systems publish to and read from. Apache Kafka.
  • Processing engines: transform large or continuous datasets on a cluster. Apache Flink, Apache Spark, and dbt inside the warehouse.
  • Orchestrators: decide when each step runs and what happens when one fails. Apache Airflow, Dagster, Prefect.
  • Source-side and telemetry pipelines: filter, reshape and route data before it crosses a network or reaches a per-GB destination. Cribl Stream, Vector, Expanso.

If you only read one line: pick the category from the job you have, then pick the tool inside it by where your data lives, who will operate it, and how it is billed. The rest of this guide covers each of those decisions, the twelve tools, and a first pipeline you can run on a laptop.

The 12 data pipeline tools at a glance

Tool Job License Pricing model Pick it when
Fivetran Managed ELT Commercial Monthly active rows You want SaaS and database data in a warehouse without running anything
Airbyte ELT ELv2 (source-available) + commercial cloud Self-hosted free; cloud by volume or capacity You want ELT connectors you can self-host or modify
Apache Kafka Event streaming Apache 2.0 Free; managed services bill by capacity Many services need the same stream of events
Apache Flink Stream processing Apache 2.0 Free; managed services vary You need stateful, event-time stream processing
Apache Spark Batch and stream processing Apache 2.0 Free; Databricks bills in DBUs You transform very large datasets on a cluster
dbt In-warehouse transformation Apache 2.0 (Core) Core free; commercial cloud Your data is already in the warehouse and you model it in SQL
Apache Airflow Orchestration Apache 2.0 Free; managed by Astronomer, MWAA, Cloud Composer You schedule many dependent jobs across tools
Dagster Orchestration Apache 2.0 Free; Dagster+ commercial You think in data assets rather than tasks
Prefect Orchestration Apache 2.0 Free; Prefect Cloud by seats and workspaces You want orchestration that is plain Python
Cribl Stream Observability pipeline Commercial Free up to 1 TB/day; credits above You route and trim logs and metrics before a SIEM
Vector Observability pipeline MPL 2.0 Free You want an open-source agent for logs and metrics
Expanso Edge-to-cloud data pipelines Commercial First 5 nodes free; $50 per node per month Data starts in many places and should be processed where it is created

Sources for every license and price are linked in each tool’s section below.

What data pipeline software actually does

Every pipeline has the same three parts: a source it reads from, processing that filters, reshapes or enriches records, and a destination it writes to. What separates the tools is where those parts run and how often.

  • ETL (extract, transform, load) transforms data before loading it. ELT loads raw data into a warehouse first and transforms it there, which is the model Fivetran, Airbyte and dbt are built around.
  • Batch pipelines run on a schedule over a bounded set of data. Streaming pipelines process records continuously as they arrive. Many teams need both.
  • Orchestration is a separate concern. An orchestrator runs other tools in order; it does not move or transform data itself.

A useful test for any product page: does this tool move bytes, transform them, or decide when something else runs? Most do one of those three well.

How to choose a data pipeline tool

Work through these in order. The first two usually narrow the field to one category.

1. The job. Getting SaaS data into a warehouse, sharing events between services, transforming large datasets, scheduling work, and trimming telemetry before it lands are five different problems. A tool built for one is usually awkward at the others.

2. Latency. If a report can be a few hours old, batch ELT is simpler and cheaper to run. If a fraud check, alert or dashboard needs data within seconds, you need a streaming tool, and you will operate it continuously.

3. Where the data starts. Most comparison articles assume your data is already in a cloud. If it starts in stores, plants, vehicles, branch offices or several cloud regions, the question becomes where processing runs. A central cluster means shipping everything to it first, paying the network cost, and only then deciding what to keep.

4. Who operates it. Open-source engines cost nothing to license and a lot of engineering time to run well. Managed services trade that time for a bill. Be honest about which of the two your team has more of.

5. How it is billed. Rows synced, gigabytes ingested, compute units, seats and nodes each grow with a different part of your business. Model next year’s volume against each pricing unit before you sign, because the cheapest tool at today’s volume is often not the cheapest at double it.

6. Connectors and extensibility. List the sources and destinations you need now and in a year, then check each one against the vendor’s connector list. For anything missing, find out how you would add it and who maintains it.

7. Control over the data. If some records must not leave a region or must lose certain fields before they are stored, check whether the tool can apply that rule before the data moves, or only after it has already landed somewhere. See our guide to data residency requirements for what the rules typically require.

Managed ingestion (ELT) tools

Fivetran

Fivetran runs connectors that copy data from SaaS applications and databases into a warehouse or lake, and handles schema changes and incremental syncs for you. Its pricing page lists 700+ fully managed connectors. It is billed by monthly active rows (MAR), and the free plan covers 500,000 MAR for connections (Fivetran pricing).

Pick it when your sources are mainstream SaaS tools and databases, your destination is a cloud warehouse, and you would rather pay than maintain extraction code. Skip it when you need to process data in flight, handle continuous event streams, or run inside networks the service cannot reach. For a wider set of options, see our Fivetran alternatives comparison.

Airbyte

Airbyte covers the same ELT job with connectors you can run yourself. Its homepage says it connects 700+ systems. The self-hosted edition, Airbyte Core, is free; Airbyte Cloud is paid (Airbyte pricing). Note the license: Airbyte’s connectors and public repositories are under the Elastic License 2.0 (ELv2) (Airbyte licenses), which is source-available rather than an OSI open-source license. That matters if you plan to offer it as a service to others.

Pick it when you want ELT connectors you can host, inspect and extend. Skip it when nobody on the team wants to operate it; then a managed service is cheaper in practice.

Event streaming

Apache Kafka

Kafka is a distributed, durable event log. Producers write events to topics, and any number of consumers read them independently. It is licensed under Apache 2.0. Since Kafka 4.0 (March 2025), it runs without ZooKeeper and uses KRaft mode by default (release announcement), which removed a whole separate system from a typical deployment.

Kafka moves and stores events. It does not transform them on its own, so it is usually paired with Kafka Streams, Flink or a consumer application. Managed Kafka exists from several vendors; Confluent Cloud, for example, bills Kafka clusters by capacity units per hour plus networking and storage (Confluent pricing).

Pick it when several services need the same events, you want them decoupled, and you can run a cluster or pay for one. Skip it when you only need to move data from A to B once; a direct pipeline is simpler.

Processing engines

Flink is a stream processor built for stateful computation over unbounded data, with exactly-once state consistency and event-time processing as core features (flink.apache.org). It is licensed under Apache 2.0.

Pick it when you need windowed aggregations, joins across streams, or results that must be correct even when events arrive late or out of order. Skip it when your team has no appetite for running a stateful cluster, or when simple per-record filtering is all you need.

Apache Spark

Spark is a distributed engine for large-scale batch processing, SQL and machine learning, licensed under Apache 2.0. Structured Streaming adds stream processing on the same SQL engine (Spark docs). Databricks, a managed Spark platform, bills in Databricks Units (DBUs) (Databricks pricing).

Pick it when you transform very large datasets, train models on them, or already run a lakehouse. Skip it when your data is small or arrives at many remote sites; a cluster is heavy machinery for either.

dbt

dbt transforms data that is already inside a warehouse. Analysts write SELECT statements, and dbt turns them into tables and views, with tests and documentation alongside (dbt on GitHub). dbt Core is licensed under Apache 2.0, and dbt Labs sells a commercial cloud product.

Pick it when your ELT tool has landed raw data in the warehouse and you want version-controlled SQL models on top. Skip it when the data never reaches a warehouse; dbt does not move data.

Orchestration tools

Orchestrators hold the dependency graph: run this after that, retry on failure, alert when last night’s run broke. They are not processing engines, and a common mistake is to put heavy transformation inside orchestrator tasks. It works at small volume and becomes the bottleneck later.

Apache Airflow

Airflow is a long-established orchestrator, licensed under Apache 2.0, with workflows written as Python DAGs. Managed versions include Astronomer, Amazon MWAA and Google Cloud Composer (Airflow ecosystem).

Pick it when you coordinate many jobs across different tools and want the largest pool of integrations and people who know it. Skip it when a single tool’s own scheduler already covers your needs.

Dagster

Dagster, licensed under Apache 2.0, models pipelines as the data assets they produce rather than as tasks, which makes lineage and freshness easier to reason about. Dagster+ is the commercial platform, with a 30-day free trial (Dagster pricing).

Pick it when you are starting fresh and care about which tables are stale and why. Skip it when you already have a working Airflow estate; migration rarely pays for itself.

Prefect

Prefect, licensed under Apache 2.0, turns ordinary Python functions into flows and tasks with decorators. Prefect Cloud prices by seats and workspaces rather than usage and has a free Hobby tier (Prefect pricing).

Pick it when your pipelines are Python code and you want the least framework around them. Skip it when you need the breadth of Airflow’s integrations.

Source-side and telemetry pipelines

These tools sit between the systems that produce logs, metrics, events and sensor readings and the platforms that store them. Their main job is to decide what is worth sending, and in what shape, before a per-GB destination bills for it. For the background on this category, see what an observability pipeline is.

Cribl Stream

Cribl Stream routes, filters and reshapes observability data between sources and destinations such as Splunk, Datadog and object storage. The free plan covers up to 1 TB/day; paid plans are billed in credits (Cribl Stream pricing). Cribl also sells a separate agent, Cribl Edge, for collection on hosts.

Pick it when your main problem is observability volume and you want a visual pipeline builder. Skip it when the data is business events or sensor telemetry rather than observability data.

Vector

Vector is an open-source agent and aggregator for observability data, licensed under MPL 2.0 and maintained by Datadog (vector.dev).

Pick it when you want a free, self-operated way to collect, transform and forward logs and metrics. Skip it when you need central management of pipelines across a large fleet without building it yourself.

Expanso

Expanso runs data pipelines on nodes you choose, close to where the data is created: a server in a store, a gateway in a plant, a VM in each cloud region, or a Kubernetes cluster. A pipeline is YAML: an input, a list of processors and an output. Expanso Edge is the runtime that executes it, as a single binary on Linux, macOS or Windows. Expanso Cloud is the managed control plane that stores pipeline configurations, assigns jobs to nodes by label, and collects their metrics and health (architecture). Pipeline records flow from sources through your nodes to the destinations you configure; the control plane receives configuration and telemetry, not the data stream.

Pick it when data starts in many places and you want to filter, reshape, enrich or route it before it crosses a network or reaches a per-GB-priced platform. The same runtime handles logs, MQTT sensor telemetry, HTTP webhooks, files and message queues such as Kafka and NATS; the component reference lists more than 50 inputs, 80 processors and 55 outputs. The same job runs the same way on a laptop, an edge box or a cloud VM.

Skip it when you need a workflow orchestrator, a batch query engine or a data warehouse. Expanso is none of those, and it works alongside them. It has no Iceberg, Delta Lake or Hudi table component yet, and its change-data-capture inputs have not completed a verified run, so for database replication into a warehouse a managed ELT tool is the safer choice today. The docs keep this list current under where Expanso is not the right answer.

Pricing: the first five nodes are free. Pro is $50 per active node per month, with no volume charges; Enterprise pricing is custom (Expanso pricing).

Run your first pipeline

This runs Expanso Edge in local mode, which needs no account, no cloud connection and no credentials. The pipeline reads a log file, drops debug lines, and tags each remaining record with the site it came from. It was run as written on Expanso Edge v2.1.21.

Install the runtime and the CLI (Linux or macOS):

curl -fsSL https://get.expanso.io/edge/install.sh | bash
curl -fsSL https://get.expanso.io/cli/install.sh | sh

In an empty directory, create a sample log file, app.log:

{"level":"debug","msg":"cache warm","ms":3}
{"level":"info","msg":"checkout ok","ms":41}
{"level":"error","msg":"payment timeout","ms":5012}
{"level":"debug","msg":"heartbeat","ms":1}
{"level":"warn","msg":"slow disk","ms":870}

Then the job, first-pipeline.yaml:

name: first-pipeline
type: pipeline
config:
  input:
    file:
      paths: ["./app.log"]
  pipeline:
    processors:
      - mapping: |
          root = if this.level == "debug" {
            deleted()
          } else {
            this.merge({"site": "store-042"})
          }
  output:
    stdout:
      codec: lines

Start the runtime in local mode from the same directory, so ./app.log resolves. --data-dir ./data keeps its state in this directory, apart from any node you have already set up on this machine:

expanso-edge run --local --data-dir ./data

In a second terminal, in the same directory, deploy the job:

export EXPANSO_CLI_ENDPOINT=http://localhost:9010
expanso-cli job deploy first-pipeline.yaml

The first terminal prints the three records that survived, each with its new site field. The two debug lines are gone. Records can print in a different order from the file:

{"level":"error","ms":5012,"msg":"payment timeout","site":"store-042"}
{"level":"info","ms":41,"msg":"checkout ok","site":"store-042"}
{"level":"warn","ms":870,"msg":"slow disk","site":"store-042"}

Stop the runtime with Ctrl+C when you are done; deleting the directory removes everything it created. To run the same job on real nodes, swap stdout for a destination from the component reference and follow the Expanso Cloud quickstart. For a larger worked example, the log reduction recipe batches, groups and archives logs end to end.

Common combinations

Most production stacks combine categories rather than pick one tool:

  • Analytics on SaaS data: Fivetran or Airbyte loads raw data into the warehouse, dbt models it, and Airflow, Dagster or Prefect schedules the runs.
  • Event-driven applications: services publish to Kafka, Flink computes aggregates and joins, and results land in a database or another topic.
  • Logs and telemetry from many sites: Expanso, Cribl or Vector filters and shapes data near the source, and only the useful part goes on to Splunk, a data lake or a warehouse. For the cost side of that pattern, see how to reduce Splunk costs.

FAQ

What are data pipeline tools?

Data pipeline tools are software that moves data from sources to destinations and processes it along the way. The category includes managed ingestion (ELT) services, event streaming platforms, processing engines, orchestrators and source-side telemetry pipelines.

What is the difference between an ETL tool and a data pipeline tool?

ETL is one kind of data pipeline: extract, transform, then load, usually into a warehouse on a schedule. “Data pipeline tool” is the broader term and also covers streaming, orchestration and pipelines that run at the source.

Is Airflow a data pipeline tool?

Airflow is an orchestrator. It schedules and monitors the steps of a pipeline and runs other tools, but it is not designed to move or transform large volumes of data itself.

Which data pipeline tools are open source?

Apache Kafka, Apache Flink, Apache Spark, Apache Airflow, Dagster, Prefect and dbt Core are licensed under Apache 2.0, and Vector under MPL 2.0. Airbyte is source-available under the Elastic License 2.0. Fivetran, Cribl Stream and Expanso are commercial, and each has a free tier.

How much do data pipeline tools cost?

Each category bills differently: ELT services by rows synced, observability pipelines by data volume, processing platforms by compute units, orchestration clouds by seats, and Expanso by node. Open-source tools have no license fee but need infrastructure and people to run them. Estimate next year’s volume against each pricing unit before comparing.

Next steps