Self-hosted observability stack cost: the real total for Prometheus, Grafana, and Loki
Running your own monitoring looks free because the software is. Adding up the instances, storage, replication, and engineer time gives a number you can actually compare against a managed bill. Here it is.
Quick answer
A production-grade self-hosted stack for a mid-sized platform (2 million active series, 100 GB/day of logs, 90-day retention) runs roughly $1,700 to $2,600 per month in infrastructure: about $900 for highly available Prometheus or Mimir compute, $200 for Loki ingesters and queriers, $180 for object storage, $120 for Grafana and Alertmanager, plus load balancers and inter-AZ transfer. The number that decides the comparison is not infrastructure, it is the 0.3 to 0.8 of an engineer needed to run it, which at loaded cost is $4,000 to $11,000 per month.
The open-source observability stack has no license fee, which makes it look free and makes the comparison against managed platforms feel like a rout. It is not free. It is a distributed database estate that you now operate, and the honest comparison needs the infrastructure bill, the object storage bill, the cross-zone transfer, and the engineering time, all in the same units as the managed quote you are holding next to it.
The reference workload
Take a concrete mid-sized platform: 250 nodes, 2 million active metric series, 100 GB per day of logs, modest tracing, 90 days of metric retention and 30 days of log retention, and a requirement that the monitoring stack survives an availability zone failure. That last requirement roughly doubles compute, which is where most back-of-envelope estimates go wrong.
Metrics tier
| Component | Sizing | Monthly cost |
|---|---|---|
| Prometheus / Mimir ingesters | 3 x r7g.xlarge (32 GB) | $470 |
| Queriers and store gateways | 2 x c7g.xlarge | $211 |
| Compactor | 1 x m7g.large | $60 |
| Local EBS gp3 for WAL/head | 3 x 300 GB | $72 |
| S3 for long-term blocks | ~3.5 TB compressed | $81 |
Two million series at roughly 2 KB of resident memory each needs about 4 GB for the head block, but real deployments need headroom for query workloads, compaction, and cardinality spikes, so 32 GB per ingester with three replicas is a normal shape. r7g.xlarge at $0.2142 per hour is $156 per month, so three is $470. The metrics tier lands around $894 per month.
Compare that to Amazon Managed Service for Prometheus for the same workload. Two million series at a 30-second scrape is 172.8 billion samples per month: $180 for the first 2 billion plus roughly $5,978 for the next 170.8 billion at $0.35 per 10 million, about $6,158 per month plus storage. Self-hosting wins the infrastructure comparison decisively at this scale, which is the honest finding and the reason people do it. See managed Prometheus pricing for the tier detail.
Logging tier
Loki's architecture is what makes self-hosted logging cheap: it indexes only labels and stores compressed log chunks in object storage, so you are not paying to build a full inverted index over every byte. At 100 GB/day with typical 8:1 compression, 30 days of retention is about 375 GB in S3, costing $9 per month. That number is not a typo. The compute to run it is the real cost.
| Loki component | Sizing | Monthly cost |
|---|---|---|
| Distributors and ingesters | 3 x m7g.large | $179 |
| Queriers | 2 x c7g.large | $105 |
| Compactor and index gateway | 1 x m7g.large | $60 |
| S3 chunks, 30 days | 375 GB | $9 |
| S3 requests | ~40M PUT/GET | ~$40 |
Roughly $393 per month for the logging tier. The managed comparison is stark: 3,000 GB per month into CloudWatch Logs at $0.50 per GB is $1,500 in ingestion alone, and Azure Monitor at roughly $2.76 per GB would be $8,280. Self-hosted logging is where the biggest raw savings live, which is why log platform comparisons usually favor it on pure infrastructure math.
Visualization, alerting, and the bits people forget
Grafana itself is cheap: two m7g.large behind a load balancer for availability, about $119 per month, plus a small managed Postgres for its database at around $30. Alertmanager runs alongside for a few dollars. Call it $160.
Then the forgotten lines. An Application Load Balancer is $16.43 per month plus LCU charges. Cross-AZ data transfer at $0.01 per GB in each direction applies to replicated writes between ingesters, and at 100 GB/day of logs with three-way replication across zones that is meaningful, roughly $60 to $120 per month. NAT gateway processing at $0.045 per GB applies if your collectors egress through one. Backups, the test environment, and a staging stack add more.
| Tier | Monthly |
|---|---|
| Metrics | $894 |
| Logs | $393 |
| Grafana and alerting | $160 |
| Load balancing and transfer | ~$150 |
| Non-production copy (~40%) | ~$640 |
| Total infrastructure | ~$2,237 |
The line that decides it
Someone has to upgrade Mimir, tune the compactor, investigate why the ingester OOMed at 3am, manage retention migrations, keep Grafana patched, and be on call for the monitoring system itself. Industry experience puts that at 0.3 of an engineer for a stable small stack and 0.8 or more once you are sharding and running multi-tenant. At a loaded cost of $220,000 per year, 0.5 of an engineer is $9,167 per month, four times the infrastructure bill.
That is the comparison that matters. Against $6,158 per month of managed Prometheus plus $1,500 of managed logs, a $2,237 infrastructure bill plus $9,167 of engineering time is $11,404 versus $7,658. Self-hosting wins on infrastructure and loses on total cost at this scale. It flips the other way as you grow, because infrastructure scales sublinearly with volume while managed per-GB pricing scales linearly: at 500 GB/day of logs the managed bill is $7,500 while the Loki tier grows to perhaps $900.
Where the crossover sits
As a rule of thumb, below roughly 50 GB/day of logs and 500,000 series, managed platforms are cheaper all in. Between there and about 300 GB/day it is genuinely close and depends on whether you already have platform engineers. Above that, self-hosting pulls away and keeps pulling. Run the numbers for your own volume rather than adopting either side's default. Every instance, volume, bucket, and load balancer in that stack is Terraform, so price the whole estate from the plan against the resource catalog and compare it against a managed quote with real numbers on both sides.
FAQ
What does a self-hosted observability stack cost to run?
For a mid-sized platform with 2 million active series, 100 GB/day of logs, and availability zone redundancy, roughly $2,237 per month in infrastructure: about $894 for the metrics tier, $393 for Loki, $160 for Grafana and alerting, $150 for load balancing and cross-AZ transfer, and about $640 for a non-production copy. That excludes engineering time, which usually exceeds the infrastructure bill.
Is self-hosted Prometheus cheaper than managed Prometheus?
On infrastructure, decisively. Two million series at a 30-second scrape produces 172.8 billion samples per month, costing about $6,158 on Amazon Managed Service for Prometheus pricing. The equivalent self-hosted tier, three r7g.xlarge ingesters, two queriers, a compactor, EBS, and S3 blocks, costs around $894 per month. The gap is real, and it is the main reason organizations at this scale self-host.
Why is self-hosted logging with Loki so cheap on storage?
Loki indexes only labels rather than building a full inverted index over log content, and stores compressed chunks in object storage. At 100 GB/day with typical 8:1 compression, 30 days of retention is about 375 GB in S3, costing roughly $9 per month plus about $40 in request charges. The compute to ingest and query it, around $344 per month, dominates. Compare that with $1,500 per month for the same volume in CloudWatch Logs ingestion alone.
How much engineer time does a self-hosted stack need?
Roughly 0.3 of an engineer for a stable small deployment and 0.8 or more once you are sharding, running multi-tenant, or handling cardinality incidents. At a loaded cost of $220,000 per year, 0.5 of an engineer is about $9,167 per month, roughly four times the infrastructure bill for a mid-sized stack. This line, not the instance costs, usually decides the build-versus-buy comparison.
At what scale does self-hosting become cheaper overall?
Below roughly 50 GB/day of logs and 500,000 series, managed platforms typically win once engineering time is counted. Between there and about 300 GB/day the comparison is close and depends on whether platform engineers already exist. Above that, self-hosting pulls ahead and keeps pulling, because infrastructure scales sublinearly with volume while managed per-GB pricing scales linearly: at 500 GB/day the managed log bill is $7,500 versus perhaps $900 for Loki.
How does C3X help compare self-hosted and managed observability?
The entire self-hosted stack is Terraform: instances, EBS volumes, S3 buckets, load balancers, and managed database for Grafana. C3X prices that estate from the plan, giving you a concrete monthly infrastructure number to set against a managed quote. That turns a build-versus-buy argument conducted in intuitions into one conducted in dollars, including the non-production copy people usually forget to count.
What to do next
Get a real number for your self-hosted stack. C3X reads your Terraform and prices every instance, volume, and bucket against a live catalog. Start with the quickstart.
Share this post
Try C3X on your own Terraform
Free and open source. No API key required. One command to install, one command to estimate.