How much does it cost to run a data pipeline processing 1 TB a day?
A streaming and batch pipeline ingesting 1 TB a day costs about 5,721 dollars a month on AWS, which is 19 cents per gigabyte. Here is the breakdown, and why the transform step is the cheapest part of it.
Quick answer
A data pipeline ingesting 1 TB a day (about 30 TB a month) through a stream, landing it in object storage, transforming it to Parquet, and serving it to analysts costs roughly 5,721 dollars a month on AWS, or about 0.19 dollars per gigabyte ingested. The largest lines are the warehouse at 1,586 dollars, stream delivery to storage at 891 dollars, raw landing storage at 707 dollars, and ad-hoc query scanning at 600 dollars. The transform itself, running Spark on Spot instances, is only 334 dollars: compute is 6 percent of the bill while movement and storage are 45 percent.
Data engineering budgets are usually discussed in terms of processing frameworks, which is the wrong end of the telescope. On a pipeline handling 1 TB a day, the Spark job that does the actual work is one of the cheapest lines. The money goes to moving bytes between services, storing them repeatedly, and letting analysts scan them.
The workload we are pricing
Assume 1 TB a day of semi-structured JSON events arriving in us-east-1, roughly 30 TB a month across about 3 billion records. Events land in a stream, are delivered to object storage in the raw zone, then a nightly Spark job converts them to partitioned Parquet at about 4:1 compression, producing 7.5 TB a month of curated data. Raw data is kept 30 days then archived; curated data is kept 90 days hot. Analysts run about 200 ad-hoc queries a day, and a BI layer serves dashboards from a small warehouse cluster.
The monthly breakdown
| Component | Specification | Monthly cost |
|---|---|---|
| Warehouse for BI | 2 x ra3.xlplus at 1.086/hr | $1,585.56 |
| Stream delivery to storage | Firehose, 30 TB at 0.029/GB | $890.88 |
| Raw zone storage | 30 TB S3 Standard at 0.023/GB | $706.56 |
| Ad-hoc queries | Athena, 120 TB scanned at 5.00 per TB | $600.00 |
| Curated zone storage | 22.5 TB of Parquet (90 days) | $529.92 |
| Stream ingestion | Kinesis 36 shards plus PUT payload units | $436.00 |
| Orchestration | Managed Airflow, small environment | $358.00 |
| Transform | EMR on Spot, 1,200 r5.2xlarge hours | $334.00 |
| Monitoring | Logs, metrics, data quality checks | $150.00 |
| Networking | VPC endpoints and residual NAT | $100.00 |
| Catalog and crawlers | Schema discovery | $30.00 |
| Total | $5,720.92 |
Query scanning is the line that varies most
Six hundred dollars of ad-hoc query cost assumes analysts scan about 20 GB per query with reasonable partition pruning. The same 6,000 monthly queries against unpartitioned JSON in the raw zone would scan perhaps 500 GB each, which is 3,000 TB, or 15,000 dollars. The difference between a well-partitioned Parquet table and a raw JSON dump is a factor of 25 on query cost alone.
| Data layout | Average scan per query | Monthly query cost |
|---|---|---|
| Raw JSON, unpartitioned | 500 GB | $15,000 |
| Parquet, unpartitioned | 125 GB | $3,750 |
| Parquet, partitioned by date | 20 GB | $600 |
| Parquet, partitioned plus column pruning | 3 GB | $90 |
This is why the transform step, at 334 dollars, is worth running even though it looks like pure overhead. Spending 334 dollars to convert JSON to partitioned Parquet saves 14,400 dollars of query scanning. It is the highest-return line in the pipeline by a wide margin.
Stream delivery costs more than stream ingestion
Firehose delivery at 0.029 per gigabyte costs 890.88 dollars for 30 TB, which is more than double the 436 dollars of stream ingestion capacity. That surprises people because delivery feels like a plumbing detail. The alternative is to consume from the stream with your own code and write batched objects yourself, which costs roughly 60 dollars of compute and removes the 891 dollar line, at the price of owning checkpointing, retries, and file sizing. For 30 TB a month that is a defensible trade; below about 5 TB a month it is not worth the engineering.
Compressing before delivery also helps directly, since Firehose bills on ingested bytes. GZIP on JSON typically achieves 5:1, so compressing at the producer would take 30 TB of billable ingest down to 6 TB and the line from 891 to about 178 dollars.
Transform framework choice
| Approach | Monthly cost for 30 TB | Notes |
|---|---|---|
| EMR on Spot | $334 | Cheapest, requires cluster management |
| EMR on-demand | $890 | Simpler, no interruption handling |
| Serverless ETL at 0.44 per DPU-hour | about $1,320 | No infrastructure, 4x the cost |
| Warehouse-native transforms | varies with cluster size | Simple if data is already loaded |
The 4x spread between Spot-based Spark and serverless ETL is real but small in absolute terms: about 1,000 dollars a month. For a team without dedicated data infrastructure engineers, that is frequently worth paying. For a team running twenty pipelines, it is 20,000 dollars a month and worth managing clusters.
Where the warehouse fits
The 1,586 dollar warehouse line is the largest single item and the most questionable. It exists to serve BI dashboards with sub-second response, which a query-on-object-storage engine cannot reliably do. If the BI workload is modest, a serverless warehouse billed per query or a well-cached semantic layer over Parquet can replace it for a few hundred dollars. If dashboards are used constantly by a large business team, the dedicated cluster earns its place. Either way it should be a deliberate decision rather than the default landing place for data.
The number to track
At 5,721 dollars for 30 TB, the pipeline costs 0.19 dollars per gigabyte ingested, or 190 dollars per terabyte. With producer-side compression, better partitioning, and a serverless warehouse, the same pipeline runs near 2,800 dollars, or 0.093 dollars per gigabyte. Track cost per gigabyte ingested month over month: it is the one metric that tells you whether the pipeline is getting more efficient or just bigger. Price the whole path from Terraform against the resource catalog before adding the next source, and see data pipeline cost optimizationfor the levers in detail.
FAQ
How much does a data pipeline processing 1 TB a day cost?
About 5,721 dollars a month on AWS for 30 TB of monthly ingest, covering stream ingestion, delivery to object storage, raw and curated storage, a Spark transform, orchestration, ad-hoc query scanning, a BI warehouse, and monitoring. That is roughly 0.19 dollars per gigabyte ingested, or 190 dollars per terabyte.
Why is the transform step the cheapest part of a data pipeline?
Because Spark on Spot instances is genuinely cheap at this volume: 1,200 r5.2xlarge hours is about 334 dollars, or 6 percent of the bill. Moving bytes between services and storing them repeatedly costs far more. But the transform has the highest return of any line, because converting JSON to partitioned Parquet saves roughly 14,400 dollars a month in query scanning.
How much does query scanning cost on a data lake?
It depends entirely on layout. At 5 dollars per terabyte scanned, 6,000 monthly queries against unpartitioned raw JSON scanning 500 GB each costs 15,000 dollars. The same queries against date-partitioned Parquet scanning 20 GB cost 600 dollars, and with column pruning down to 3 GB, only 90 dollars. Layout is a 100x lever on query cost.
Is Firehose delivery worth the cost?
At 0.029 per gigabyte, delivering 30 TB a month costs 891 dollars, more than double the stream ingestion capacity itself. Writing batched objects from your own consumer costs roughly 60 dollars of compute but means owning checkpointing, retries, and file sizing. Above about 30 TB a month that trade is defensible; below 5 TB it is not worth the engineering. Compressing at the producer cuts the line to about 178 dollars either way.
Should I use serverless ETL or managed Spark clusters?
Serverless ETL at 0.44 per DPU-hour costs about 1,320 dollars for this workload against 334 dollars for Spark on Spot, a 4x spread but only about 1,000 dollars in absolute terms. For a team without dedicated data infrastructure engineers, paying the premium is usually correct. For a team running twenty pipelines, the same ratio is 20,000 dollars a month and cluster management pays for itself.
How does C3X help control data pipeline costs?
C3X prices infrastructure from Terraform against a live catalog, so stream shards, delivery streams, storage buckets and their lifecycle rules, cluster sizes, and warehouse nodes are costed before they are applied. Since pipeline cost is dominated by movement and storage rather than compute, seeing those resources priced at design time is what keeps the per-gigabyte number from drifting upward as sources are added.
What to do next
Know your cost per gigabyte before you add the next source. C3X reads your Terraform and prices your resources against a live catalog. Start with the quickstart.
Share this post
Try C3X on your own Terraform
Free and open source. No API key required. One command to install, one command to estimate.