Every engineering team starts the same way. You pick a monitoring tool, point your logs and metrics at it, and move on. Datadog, Splunk, Elastic, New Relic, whatever the team already knows. For a while, it works.
Then the bill arrives. Your telemetry volume doubled because someone enabled debug logging in production and forgot to turn it off. Your Kubernetes cluster scaled to 200 pods, each emitting verbose health checks every 5 seconds. Your cloud provider added audit logs for a new service, and nobody noticed until the monthly invoice.
By the time most organizations notice the problem, they are spending $500K to $3M+ per year on observability platforms, and 60-80% of the data hitting those platforms is noise that nobody queries. The data observability and monitoring challenge is not collecting data. It is controlling what happens to data between collection and storage.
That is what an observability pipeline does. It sits between your data sources and your observability platforms, giving you a place to filter, transform, route, and control telemetry before it costs you money.
This post explains how observability pipelines work, what the architecture looks like, which tools exist, and how to decide whether your organization needs one.
What Is an Observability Pipeline?
An observability pipeline is a processing layer between telemetry sources and observability backends. It receives logs, metrics, and traces from your infrastructure and applications, applies transformations and routing rules, and forwards the processed data to one or more destinations.
Without a pipeline, the architecture is direct:
Sources -> Observability Platform
Every byte generated by every application, container, server, and cloud service goes straight to your monitoring tool. You pay for all of it. You search through all of it. You store all of it.
With a pipeline, the architecture becomes:
Sources -> Observability Pipeline -> Observability Platform(s)
The pipeline gives you a control point. Before data reaches any backend, you can:
Filter. Drop events that provide no analytical value. Debug logs in production, health check responses, duplicate heartbeats, verbose HTTP access logs from load balancers. If nobody queries it, stop paying to store it.
Transform. Reshape events to reduce their size or improve their usefulness. Strip unnecessary fields, parse unstructured text into structured JSON, normalize timestamps across time zones, enrich events with metadata from external sources.
Route. Send different data types to different backends. Security events go to your SIEM. Application metrics go to Prometheus. Raw logs go to S3 for cold storage at $0.023/GB/month instead of hot observability storage at $150+/GB/day. One data stream, multiple destinations, each receiving only what it needs.
Sample. High-volume, low-value streams can be sampled instead of dropped entirely. Keep 10% of debug logs for troubleshooting coverage without paying for the full 100%.
Redact. Strip PII, credentials, and sensitive fields before data leaves a boundary. Data that never reaches your observability platform never needs to be governed there. This is critical for organizations with strict data governance requirements.
The result is that your observability platforms receive less data, better data, and the right data. Costs go down. Query performance goes up. Signal-to-noise ratio improves.
Why Observability Pipelines Exist
The observability pipeline category exists because of a structural problem in how telemetry data flows through modern infrastructure.
The Volume Problem
Telemetry volume grows faster than budgets. Five years ago, a typical enterprise generated 200-500 GB/day of machine data. Today, the same enterprise generates 2-10 TB/day. Kubernetes, microservices, serverless functions, cloud audit trails, IoT sensors, and edge devices all contribute. Every new workload adds to the meter.
Observability platforms charge by volume. More data means higher costs. But the relationship between data volume and analytical value is not linear. Most of the growth comes from verbose, repetitive, low-value telemetry that nobody uses for troubleshooting, security investigation, or capacity planning.
The Vendor Lock-in Problem
Without a pipeline, your data format and routing are determined by your observability vendor. Switching from Splunk to Elastic means re-instrumenting your entire collection infrastructure. Running both simultaneously means duplicating data and paying twice.
A pipeline decouples collection from storage. Your agents and instrumentation send data to the pipeline. The pipeline sends data to whichever backends you choose. Swapping a backend or adding a new one is a configuration change, not a re-instrumentation project. This flexibility is especially valuable for data integration across multiple platforms.
The Compliance Problem
Regulations like GDPR, HIPAA, and PCI-DSS restrict where certain data can be stored and who can access it. Without a pipeline, sensitive data flows directly from source to backend. If your backend is in a different region or operated by a third party, you may have a compliance exposure.
A pipeline lets you redact, mask, or route sensitive data before it leaves a boundary. PII gets stripped at the source. Region-specific data stays in the region. Compliance requirements are enforced in the data path, not after the fact.
Observability Pipeline Architecture
Every observability pipeline follows a three-stage pattern: collect, process, route. The differences between tools are in where each stage runs and how much control you get at each step.
Stage 1: Collection
Data enters the pipeline from sources. These include:
- Application logs from stdout, files, or logging libraries
- Infrastructure metrics from system agents, cloud APIs, and orchestrators
- Distributed traces from instrumented services via OpenTelemetry, Jaeger, or Zipkin
- Cloud platform logs from AWS CloudWatch, Azure Monitor, GCP Cloud Logging
- Network telemetry from flow logs, packet captures, and DNS queries
- Custom events from message queues, databases, and business applications
Collection happens via agents running on source machines, API integrations pulling from cloud services, or direct protocol receivers (syslog, HTTP, gRPC, Kafka consumers).
The collection tier determines your pipeline’s reach. If collection only covers servers in your data center, you miss cloud-native telemetry. If it only covers cloud, you miss on-premises and edge data. The best pipelines support a broad set of inputs. The Expanso product connects to 61 input types, covering everything from syslog and file tailing to Kafka, MQTT, and cloud-native APIs.
Stage 2: Processing
This is where the pipeline earns its keep. Processing includes:
Parsing. Convert unstructured log lines into structured fields. A raw Apache access log becomes a JSON object with method, path, status code, response time, and client IP as discrete, queryable fields.
Filtering. Evaluate each event against rules and drop those that match exclusion criteria. Drop all events where log_level == "DEBUG" and environment == "production". Drop health check responses from load balancers. Drop Kubernetes liveness probe logs.
Sampling. For high-volume streams where complete filtering is too aggressive, keep a percentage. Sample verbose HTTP access logs at 10%. Sample trace spans from low-priority services at 5%. Keep 100% of error-level events.
Enrichment. Add context that isn’t present in the raw event. Look up a host’s business unit from a CMDB. Add geographic information based on IP address. Tag events with the deployment version that generated them.
Aggregation. Collapse high-cardinality data into summaries. Instead of storing every individual HTTP request, store per-minute counts by status code and endpoint. This is particularly effective for metrics pipelines where raw data points are less useful than statistical summaries.
Redaction. Apply regex patterns to strip credit card numbers, Social Security numbers, email addresses, or API keys from event payloads. Replace matched patterns with placeholder values.
Format conversion. Convert between data formats. JSON to Avro for compact storage. CSV to JSON for structured querying. Protobuf to JSON for human readability. Syslog to structured JSON for modern backends.
Stage 3: Routing
Processed data gets sent to one or more destinations. A single event can be routed to multiple backends simultaneously:
- Security events to Splunk or your SIEM
- Application metrics to Prometheus or Datadog
- Raw logs to S3 or Azure Blob for long-term archive
- Aggregated metrics to InfluxDB or ClickHouse
- Alerts to PagerDuty or Slack via webhooks
- Compliance-relevant events to a dedicated audit store
Routing decisions can be content-based. Events matching a security detection rule go to the SIEM. Events from a specific application go to that team’s preferred tool. Events above a certain severity always go to hot storage; everything else goes to cold.
This multi-destination routing is what breaks vendor lock-in. Your pipeline owns the routing logic, not your backend vendor.
Centralized vs. Edge Pipeline Architectures
Observability pipelines broadly fall into two architectural categories. Where you process data determines your cost profile, latency characteristics, and operational complexity.
Centralized Pipelines
A centralized pipeline runs as a cluster in your data center or cloud VPC. All telemetry flows from sources to the central cluster, gets processed there, and then routes to destinations.
Sources -> (network) -> Central Pipeline Cluster -> Destinations
Advantages:
- Easier to manage. One cluster, one set of rules, one deployment.
- Full visibility into all data at a single point.
- Simpler to debug processing rules because everything flows through one place.
Disadvantages:
- All raw data still travels over the network. If you generate 10 TB/day across thousands of sources, that 10 TB hits your network before any filtering happens.
- The central cluster becomes a scaling bottleneck. More data means more cluster capacity.
- Single point of failure. If the pipeline cluster goes down, telemetry stops flowing.
- Cannot reach true edge locations (cell towers, retail stores, vehicles) where network connectivity is expensive or unreliable.
Most commercial observability pipelines use this model. Cribl Stream, Mezmo Pipeline, and similar tools run as centralized processing clusters.
Edge Pipelines (Upstream Data Control)
An edge pipeline runs lightweight agents directly on the source machines. Filtering, transformation, and routing happen before data leaves the source.
Sources (with local agent) -> (filtered data over network) -> Destinations
Advantages:
- Only processed, filtered data traverses the network. If you filter 70% at the source, your network carries 70% less data.
- No centralized bottleneck. Each agent handles its own source’s data.
- Reaches anywhere you can deploy a binary: data centers, cloud VMs, edge devices, IoT gateways, telecom infrastructure.
- Resilient to network interruptions. Agents can buffer locally when connectivity drops.
Disadvantages:
- More agents to manage. Thousands of sources means thousands of agents.
- Requires fleet management to push configuration changes at scale.
- Correlation across sources requires additional architecture (the agents operate independently).
Expanso’s upstream data control plane uses this model. A single agent binary runs on each source node; its memory and CPU use depend on the pipeline it runs. It connects to 50 input types, processes data through 80 processors, and routes to 55 output destinations. Fleet management pushes configuration changes to 8,500+ nodes in under 30 seconds.
Hybrid Approach
Many organizations combine both. Edge agents handle first-pass filtering and routing at the source. A centralized pipeline handles cross-source correlation, complex enrichment that requires external lookups, and final routing logic.
This layered approach is common in financial services environments where both cost control and compliance processing are critical.
Observability Pipeline Tools
The observability pipeline market includes open-source projects, commercial platforms, and vendor-specific tools. Here is an honest assessment of each category.
OpenTelemetry Collector
OpenTelemetry (OTel) is the CNCF standard for telemetry collection and processing. The OTel Collector is an open-source agent and gateway that receives, processes, and exports telemetry data.
Strengths: Vendor-neutral. Strong trace and metric support. Large ecosystem of receivers, processors, and exporters. Growing community.
Limitations: Primarily designed for traces and metrics. Log support is maturing but less battle-tested. Configuration can be complex for advanced routing. No built-in fleet management for large deployments. Requires you to build and maintain your own deployment infrastructure.
Best for: Organizations standardizing on OpenTelemetry instrumentation who want a vendor-neutral collection layer.
Cribl Stream
Cribl Stream is a commercial observability pipeline focused on log routing and transformation. It runs as a centralized cluster and provides a visual interface for building processing rules.
Strengths: Strong log processing capabilities. Visual pipeline builder. Good Splunk integration. Reduces Splunk ingest costs effectively.
Limitations: Centralized architecture means all raw data still traverses the network. Pricing scales with throughput, which can get expensive at high volumes. Less suited for edge deployments where you cannot run a cluster. Primarily log-focused. Metrics and traces are secondary.
Best for: Organizations with a large Splunk investment looking to reduce ingest costs without changing their architecture.
Vector (by Datadog)
Vector is an open-source observability data pipeline written in Rust. It supports logs, metrics, and traces with a focus on performance and reliability.
Strengths: High performance. Low resource consumption. Open source with strong community. Supports both agent and aggregator deployment modes. VRL (Vector Remap Language) is powerful for transformations.
Limitations: Now owned by Datadog, which raises vendor-neutrality questions. Fleet management is not built in. Community support only for the open-source version.
Best for: Performance-sensitive deployments where an open-source, lightweight pipeline agent is preferred.
Fluent Bit / Fluentd
Fluentd and its lightweight counterpart Fluent Bit are CNCF projects for log collection and forwarding. Fluent Bit is widely used as a Kubernetes log collector.
Strengths: Lightweight (especially Fluent Bit). Strong Kubernetes integration. Large plugin ecosystem. Mature and battle-tested.
Limitations: Primarily log-focused. Transformation capabilities are more limited than dedicated pipeline tools. Configuration via text files can be cumbersome at scale. No built-in fleet management.
Best for: Kubernetes-native log collection where simplicity and low resource usage matter.
Splunk Edge Processor
Splunk’s own pre-ingest processing tool. Sits between data sources and Splunk Cloud to filter, mask, and route events before indexing.
Strengths: Integrated with Splunk Cloud. Uses SPL2, familiar to Splunk teams. Included with Splunk Cloud license.
Limitations: Only works with Splunk Cloud (not Splunk Enterprise). Only routes to Splunk destinations. Not a general-purpose pipeline. See our detailed Splunk Edge Processor comparison for a deeper analysis.
Best for: Splunk Cloud-only environments with straightforward filtering needs.
Expanso (Upstream Data Control)
Expanso takes the edge pipeline approach. Lightweight agents run on source machines and handle filtering, transformation, and multi-destination routing before data leaves the source.
Strengths: Processes data at the source, eliminating network transfer of unwanted data. Supports 50 inputs, 80 processors, 55 outputs. Fleet management at 8,500+ node scale (tested at 100,000+). A single binary runs on any platform. Backend-agnostic routing.
Limitations: Requires deploying agents to source machines. Edge-first model means cross-source correlation requires additional architecture.
Best for: Multi-backend environments, edge deployments, organizations generating 1+ TB/day that need to control costs across multiple observability platforms.
When You Need an Observability Pipeline
Not every organization needs an observability pipeline. If you have a few dozen servers, one monitoring tool, and a telemetry bill under $50K/year, direct collection is probably fine.
Here are the signals that you need one:
Your Observability Bill Exceeds $200K/Year
At higher spend levels, modest changes in billable telemetry volume can materially affect cost. Build the business case from your own ingest, retention, query, and contract data; then validate any projected reduction with a representative pilot before committing to a rollout.
You Send Data to Multiple Backends
If logs go to Splunk, metrics go to Datadog, and traces go to Jaeger, you have three separate collection configurations to maintain. A pipeline centralizes collection and handles routing, reducing operational complexity. Adding a new backend becomes a routing rule change, not a re-instrumentation project.
Most of Your Data Is Noise
Run this test: look at your top 10 sourcetypes or log groups by volume. For each one, ask your team when they last queried it. If more than half have no recent queries, you are paying to store data nobody uses. A pipeline filters that noise before it reaches your platform.
You Have Compliance Requirements
GDPR, HIPAA, PCI-DSS, and similar regulations require control over where data goes and what it contains. A pipeline enforces redaction and routing rules in the data path, turning compliance from a manual audit exercise into an automated enforcement layer.
You Run Edge or Distributed Infrastructure
Cell towers, retail locations, oil rigs, hospital devices, vehicles, manufacturing floors. If your data sources are geographically distributed and connected by expensive or unreliable networks, you need processing at the source. A centralized pipeline cannot help if the problem is data volume on the network between source and pipeline. This is especially relevant for energy sector deployments with remote infrastructure.
Your Query Performance Is Degrading
When an observability platform processes less noise, some queries may run faster. If searches or dashboards are slow, measure whether low-value data is competing for resources before assuming the platform is the cause, and compare query latency before and after a controlled filtering test.
How to Evaluate an Observability Pipeline
If you have decided you need a pipeline, here is what to evaluate:
Data Type Coverage
Does the pipeline handle your data types? Logs are table stakes. Metrics, traces, and structured events matter too. If you have IoT sensor data, binary payloads, or database CDC streams, make sure the pipeline supports them natively.
Deployment Model
Centralized cluster or edge agents? The answer depends on your architecture. If your sources are in a few cloud regions, a centralized pipeline may be fine. If your sources span thousands of locations with limited network connectivity, you need edge agents.
Transformation Language
How powerful is the processing engine? Can it parse unstructured logs? Convert between formats? Apply conditional logic? Sample selectively? Every pipeline has a transformation language, SPL2 for Splunk Edge Processor, VRL for Vector, Bloblang for Expanso. Try writing your top 5 transformation use cases in each language and see which one fits.
Fleet Management
If you deploy agents to hundreds or thousands of nodes, how do you update their configuration? Manually SSH-ing into machines does not scale. Look for centralized configuration management with atomic rollouts, health monitoring, and rollback capabilities.
Output Destinations
Count the destinations you use today and the ones you might use in the next two years. Make sure the pipeline supports them. A pipeline that only routes to one vendor’s platform is not a pipeline: it is a vendor tool.
Total Cost of Ownership
Compare the pipeline cost against projected savings. Include infrastructure costs for running the pipeline, licensing fees, and the operational cost of managing it. Use customer-specific contract and workload data instead of assuming a fixed payback period or return.
Common Mistakes When Deploying Observability Pipelines
Filtering Too Aggressively
The first instinct is to filter everything. Teams set up rules, see their bill drop, and celebrate. Then an incident happens and the logs they need were filtered out.
Start with data profiling. Identify what your team actually queries. Filter conservatively at first. Sample high-volume streams instead of dropping them entirely. Route filtered data to cheap cold storage rather than discarding it forever. You can always tighten filters later.
Ignoring the Network Tier
A centralized pipeline still requires all raw data to cross the network. If your problem is network bandwidth between edge sites and your data center, a centralized pipeline does not help. Processing must happen upstream of the bottleneck, not downstream of it.
Treating the Pipeline as Set-and-Forget
Your infrastructure changes. New applications get deployed. Existing applications change their log formats. Telemetry volume shifts. A pipeline that was tuned six months ago may be passing through noise it should filter, or filtering data that has become valuable.
Review pipeline rules quarterly. Monitor filter hit rates. If a filter rule never matches, it is either stale or your data changed. If a new sourcetype is passing through unfiltered, it needs a rule.
Building Instead of Buying
Some teams build their own pipeline using Kafka, custom consumers, and scripts. This works at small scale. At 1+ TB/day with dozens of sourcetypes and multiple destinations, a custom pipeline becomes a full-time engineering project. Evaluate commercial and open-source options before committing engineering resources to building from scratch. The log processing at scale challenge is well-understood, and existing tools handle it reliably.
If your telemetry costs are climbing and your data is distributed across cloud, on-prem, and edge, start with a consultation to map your data flows and identify where a pipeline would have the highest impact.
Related Articles
- Splunk Pricing in 2026: The Real Cost and How to Control It
- Splunk Architecture Explained: Components, Data Flow, and Optimization
- Splunk Edge Processor vs Upstream Data Control
- How to Reduce Splunk Costs Without Losing Visibility
- SIEM Cost Comparison: Why Your Bill Keeps Growing
FAQ
What is the difference between an observability pipeline and a log aggregator?
A log aggregator collects logs from multiple sources and centralizes them in one place. Fluentd and Logstash are examples. They focus on collection and forwarding, with limited transformation capabilities. An observability pipeline adds processing, routing, and filtering between collection and storage. It handles logs, metrics, and traces. It routes to multiple destinations. It reduces volume before data reaches your backends. A log aggregator is one component of a pipeline, not the whole thing.
How much can an observability pipeline reduce my telemetry costs?
There is no universal cost-reduction percentage for an observability pipeline. Establish the current source volume, retained event classes, destination pricing, and operational requirements; then compare those measurements with a representative upstream-filtering evaluation. Savings depend on the source data, selected rules, information that must be retained, and the downstream platform contract.
Does an observability pipeline add latency to my monitoring?
Latency depends on the input, processors, buffering, output, hardware, and network path. Measure end-to-end event timing under representative load, backpressure, restart, and destination-failure conditions before accepting a pipeline for production.
Can I use OpenTelemetry as my observability pipeline?
OpenTelemetry Collector can serve as a lightweight pipeline for traces and metrics. It supports receivers, processors, and exporters in a configurable chain. For log-heavy environments or advanced routing needs (content-based routing, format conversion, complex enrichment), you may need a more full-featured pipeline tool. Many organizations use OTel Collector for trace and metric collection and a separate pipeline for log processing and multi-destination routing.
How is an observability pipeline different from Cribl?
Cribl Stream is one implementation of the observability pipeline pattern. It runs as a centralized processing cluster and focuses on log routing and transformation, with strong Splunk integration. An observability pipeline is the broader architectural concept. Other implementations include edge-based agents (Expanso), open-source tools (Vector, Fluent Bit), and vendor-specific processors (Splunk Edge Processor). The right implementation depends on your deployment model, data types, and destination requirements.
Do I need an observability pipeline if I only use one monitoring tool?
Maybe. Even with a single backend, a pipeline can reduce costs by filtering noise before ingest, improve compliance by redacting sensitive data in transit, and improve query performance by ensuring only high-value data reaches your platform. The ROI is smaller than in multi-backend environments, but if your single-tool bill exceeds $200K/year, the math usually works. If you use Splunk specifically, see our Splunk pricing guide for targeted cost reduction strategies.
What data should I never filter from an observability pipeline?
Never filter security events that your SOC actively investigates, audit logs required for compliance, error and fatal-level application logs, and authentication events. These data types have high analytical value relative to their volume. Focus filtering on debug logs, health checks, verbose HTTP access logs, duplicate events, and infrastructure noise that nobody queries. When in doubt, route to cheap cold storage instead of dropping entirely.
How do I get started with an observability pipeline?
Start with data profiling. Identify your top 10 sourcetypes by volume and check which ones your team actually queries. The gap between “what you collect” and “what you use” is your filtering opportunity. Then evaluate tools based on your architecture (centralized vs. edge), data types, and destination requirements. Most teams start with a single high-volume, low-value data source as a pilot, demonstrate ROI, and expand from there. Talk to the Expanso team if you want help scoping the opportunity.
