Trace sampling economics: what 1% sampling actually costs you
Sampling is the main cost lever in distributed tracing, and picking a rate is a trade between spend and the odds of capturing the rare failure you needed. Here is the math on both sides.
Quick answer
Trace cost scales linearly with sampling rate, so 1% sampling costs 1% of what 100% would. AWS X-Ray charges $5.00 per million traces recorded, so 1 billion requests per month at 100% sampling costs $5,000, at 10% costs $500, and at 1% costs $50. The catch is statistical: at 1% sampling an error affecting 0.1% of requests yields about 1,000 sampled traces from a billion requests, which is plenty, but a bug hitting 10 requests per month is captured with 10% probability. The fix is tail sampling: keep 100% of errors and slow traces, sample the boring successes at 0.1%.
Distributed tracing is the most expensive observability signal per unit of insight, and sampling is the only real lever. Every platform prices it the same way: you pay per trace or per span ingested, the rate multiplies directly into the bill, and a request volume that is fine at 1% is ruinous at 100%. The interesting question is not whether to sample but what the sampled-away traces were worth.
What traces cost
| Backend | Unit price | 1B requests/month at 100% |
|---|---|---|
| AWS X-Ray (recorded) | $5.00 per million traces | $5,000 |
| AWS X-Ray (retrieved/scanned) | $0.50 per million traces | varies by query |
| Google Cloud Trace | $0.20 per million spans after 2.5M free | $2,000 at 10 spans/trace |
| Typical SaaS APM | $1.70 to $2.50 per million spans | $17,000 to $25,000 at 10 spans/trace |
| Self-hosted on S3 + compute | storage plus query compute | see below |
Note the span multiplier. A single trace through an API gateway, three services, a cache, and a database can easily produce 10 to 30 spans. Backends that price per span rather than per trace multiply your volume by that factor, so a service mesh that adds two spans per hop can double the bill without adding a single new service. Count spans, not requests, when you forecast.
The linear part
Sampling arithmetic is mercifully simple. At 1 billion requests per month into X-Ray, 100% sampling costs $5,000, 10% costs $500, 5% costs $250, and 1% costs $50. The X-Ray default sampling rule of one request per second per host plus 5% of additional requests exists precisely because someone at AWS did this arithmetic. If your APM bill is dominated by traces and your sampling rate is above 10%, you have found your optimization in one setting change.
The statistical part
The cost of sampling is captured traces you do not have when you need them. Model it as a binomial: if a failure occurs N times in the sampling window and you sample at rate p, the probability you captured at least one is 1 minus (1 minus p) to the power N.
| Failures per month | p = 1% | p = 10% | p = 50% |
|---|---|---|---|
| 10 | 9.6% | 65% | 99.9% |
| 100 | 63% | 99.997% | ~100% |
| 1,000 | 99.996% | ~100% | ~100% |
| 100,000 | ~100% | ~100% | ~100% |
The table says something useful: for anything happening more than about 500 times a month, 1% sampling is statistically fine, and you are paying 100x for traces of common events you could have characterized from a hundredth of the data. The problem is entirely at the top of the table, the rare failures, and those are exactly the ones worth debugging.
Tail sampling changes the trade
Head sampling decides at the start of a request, before anything has gone wrong, so it treats a 500 response and a healthy 200 identically. Tail sampling buffers the complete trace and decides after, which means you can keep every error and every trace over a latency threshold while sampling successful fast requests at a fraction of a percent.
The economics are excellent. Suppose 1 billion requests per month with a 0.2% error rate and a further 0.5% exceeding your latency SLO. Keeping 100% of those 7 million interesting traces plus 0.1% of the remaining 993 million gives about 7.99 million traces, costing roughly $40 on X-Ray pricing versus $5,000 for full sampling. You retained every failure and paid 0.8% of the price.
The cost moves to the collector. Tail sampling requires buffering all spans of a trace until it completes, which means the OpenTelemetry Collector needs memory proportional to in-flight traces times span size, and all spans of a trace must reach the same collector instance, requiring a load-balancing exporter layer. Budget a collector tier of two or three c7g.xlarge instances at roughly $0.1445 per hour each, about $316 per month for three, which is still trivially cheaper than the $4,960 you saved. The APM cost picture usually improves sharply once this is in place.
Retention and retrieval, the forgotten meters
X-Ray charges $0.50 per million traces retrieved or scanned, separately from the $5.00 recording charge. Aggressive dashboards and automated trace analysis can make retrieval a material line of its own. Most SaaS backends bundle 15 or 30 days of trace retention and charge extra beyond it, at which point exporting completed traces to S3 in Parquet at $0.023 per GB-month and querying with Athena at $5.00 per TB scanned becomes a reasonable cold tier for anything older than a month.
Setting the policy
A defensible default: tail sampling with 100% of errors, 100% of traces exceeding the p99 latency target, 100% of a small set of business-critical endpoints, and 0.1% to 1% of everything else, with per-service overrides so a team debugging a specific service can raise its rate temporarily. Make the rate a configuration value, not a code constant, so raising it during an incident does not need a deploy.
The collector fleet, its autoscaling group, and the storage behind your traces are all Terraform. Price them from the plan against the resource catalog so the cost of the tracing pipeline is known before the traces start flowing through it.
FAQ
How much do distributed traces cost?
AWS X-Ray charges $5.00 per million traces recorded plus $0.50 per million retrieved or scanned. Google Cloud Trace charges $0.20 per million spans after 2.5 million free. SaaS APM products typically land between $1.70 and $2.50 per million spans. The span multiplier matters: one request through a gateway, three services, a cache, and a database can produce 10 to 30 spans, so per-span pricing multiplies your request volume accordingly.
What does 1% trace sampling actually give up?
For common events, almost nothing. A failure occurring 1,000 times per month is captured with 99.996% probability at 1% sampling. The gap is at the rare end: a bug hitting 10 times per month is captured with only 9.6% probability at 1%, rising to 65% at 10% sampling. So 1% sampling is statistically sound for anything happening more than roughly 500 times per month and unreliable for genuinely rare failures, which are often the ones worth debugging.
How does tail sampling improve trace economics?
Tail sampling decides after the trace completes, so it can keep 100% of errors and slow traces while sampling routine fast successes at 0.1%. On 1 billion requests per month with a 0.2% error rate and 0.5% exceeding the latency SLO, keeping all 7 million interesting traces plus 0.1% of the rest yields about 7.99 million traces, roughly $40 on X-Ray pricing versus $5,000 for full sampling, with no failures lost.
What does tail sampling cost to run?
It moves cost to the collector tier. Tail sampling buffers every span of a trace until completion, so the OpenTelemetry Collector needs memory proportional to in-flight traces, and all spans of a trace must reach the same instance, requiring a load-balancing exporter layer. Three c7g.xlarge instances at roughly $0.1445 per hour cost about $316 per month, which is a fraction of the several thousand dollars saved on ingestion.
Should I keep traces longer than 30 days?
Rarely in the hot tier. Most backends bundle 15 or 30 days and charge extra beyond, and interactive trace debugging is almost always about the recent past. For longer retention, export completed traces to S3 in Parquet at $0.023 per GB-month and query with Athena at $5.00 per TB scanned. That gives you historical analysis at a small fraction of hot-tier retention pricing, with query latency measured in seconds rather than milliseconds.
How does C3X help with tracing cost?
The collector fleet, its autoscaling group, the load balancer in front of it, and the storage behind traces are all Terraform resources. C3X prices them from the plan, so the cost of standing up a tail-sampling collector tier or expanding trace storage is visible in the pull request. That lets you compare the collector cost against the ingestion savings before committing to the architecture.
What to do next
Price your tracing pipeline before the spans start flowing. C3X reads your Terraform and prices collectors and storage against a live catalog. Start with the quickstart.
Share this post
Try C3X on your own Terraform
Free and open source. No API key required. One command to install, one command to estimate.