Notes from the team building C3X. Cost estimation for Terraform, the economics of cloud infrastructure, and how to ship FinOps tooling without a SaaS gate.
The three big managed Kubernetes services charge similar control-plane fees but diverge sharply on node pricing, load balancers, and free tiers. A like-for-like cluster can differ by 15 to 20 percent across them. Here is the comparison with real numbers.
Before a single pod schedules, most managed Kubernetes clusters bill a flat hourly fee for the control plane. It is small per cluster and enormous across a fleet of fifty. Here is what each provider charges and how to stop paying it fifty times.
Node size is one of the highest-leverage cost decisions in Kubernetes and most teams pick by habit. Too small and system overhead eats your capacity, too big and a single pod strands half a node. Here is how to size node pools with real numbers.
Most Kubernetes clusters run at 25 to 40 percent CPU utilization, which means well over half the node bill buys nothing. Bin packing is the discipline of fitting pods onto fewer nodes. Here is how the scheduler decides and what to change.
Across most clusters, pods request roughly three times the CPU they actually consume, and the cluster buys nodes for the requests. That gap is the largest single line of avoidable Kubernetes spend. Here is where it comes from and how to close it.
PersistentVolumeClaims quietly provision real cloud disks, and those disks keep billing long after the pod is gone. Storage class choice, over-sized claims, and orphaned volumes routinely add 15 to 25 percent to a cluster bill. Here are the numbers.
Every Service of type LoadBalancer provisions a real cloud load balancer with its own hourly fee. Teams that expose each microservice this way pay hundreds per month for routing that one ingress controller would handle. Here is the math and the fix.
Horizontal and Vertical Pod Autoscalers solve different problems and have opposite cost behaviours. HPA can raise your bill while improving latency, VPA usually lowers it by shrinking requests. Here is how each affects the node count you pay for.
GPU nodes cost 10 to 40 times a general-purpose node per hour, and Kubernetes gives you a whole GPU per pod by default. Idle GPU nodes are the most expensive waste in any cluster. Here are the rates and the sharing strategies.
Shared clusters are far cheaper than cluster-per-team, but only if one team cannot consume the capacity everyone else paid for. ResourceQuotas, LimitRanges, and priority classes are the controls that make shared infrastructure fair. Here is how to set them.
Every sidecar container multiplies across every pod in the cluster. A 100m CPU and 128 MB sidecar on 800 pods is 80 vCPU and 100 GB of pure infrastructure. Here is how to measure the tax and decide which sidecars earn their keep.
In-cluster cost agents need deployment, upgrades, RBAC, and their own compute, and they only report after money is spent. Reading cost straight from Terraform gives you the number before the cluster exists. Here is the tradeoff.
Clusters multiply one team at a time until a fleet of twenty exists with no one who chose it. Each carries a control plane fee, a system-pod floor, and its own idle headroom. Consolidation routinely removes 50 to 70 percent of fixed cluster cost.
Non-production clusters run 168 hours a week and are used for about 45. Scaling node pools to zero outside working hours is the least controversial saving in Kubernetes. Here is what it saves and what breaks.
Clusters scale up reliably and scale down almost never, which is why node counts only ever ratchet upward. The causes are a short list of blockers, most of them fixable in an afternoon. Here is the diagnosis and the tuning.
Spot nodes cost 60 to 90 percent less but can disappear with two minutes of warning. The question is not whether to use them but what percentage of the cluster they should be. Here is how to choose the ratio by workload.
Spreading pods across availability zones is good for resilience and expensive for chatty microservices. Every cross-zone hop is billed in both directions. Here is how much it adds and how topology aware routing removes most of it.
Batch jobs and CronJobs look free because they finish. In practice they hold nodes open, block scale-down, and provision capacity for a peak that lasts four minutes an hour. Here is how to make batch nearly free.
An H100 costs roughly three times an A100 per GPU-hour, but it can be two to four times faster on modern training workloads. The cheaper GPU is the one with the lower cost per unit of work, not the lower hourly rate. Here is how to do that math.
The T4 has been the default cheap inference GPU for years, but the L4 costs about 53 percent more per hour while delivering two to three times the throughput on modern models. For inference fleets, that inversion is worth real money.
A hosted API charges per token with no floor. A self-hosted model charges per GPU-hour whether you use it or not. The crossover sits at a specific monthly token volume, and most teams guess it wrong by an order of magnitude.
Adding pgvector to a Postgres instance you already run looks free. A managed vector service starts at several hundred dollars a month. The honest comparison involves index memory, replica sizing, and how much your team wants to operate.
Fine-tuning cost is a product of four numbers: model size, dataset tokens, GPU rate, and how many times you will redo it. Get those on paper and a run that felt unbounded turns into a figure you can approve.
Spot GPUs cost 50 to 70 percent less and can vanish with two minutes of warning. Whether that trade is good depends on one number: how much work you lose per interruption, which is a function of your checkpoint interval.