data platform
14 articles on data platform — what drives the cost, how it is priced, and where the savings actually are.
Snowflake credit pricing explained: what a credit actually costs you
Snowflake bills compute in credits, storage in terabytes, and a handful of serverless features on their own meters. Credits look cheap until you multiply by warehouse size and hours. Here is how the math really works.
Snowflake warehouse sizing cost: when bigger is cheaper and when it is not
Every warehouse size step doubles the credit burn. Sometimes it halves the runtime and the cost stays flat, sometimes it doubles your bill for nothing. Here is how to tell the two apart before you resize.
Data warehouse auto suspend: the setting that decides half your bill
Idle warehouse time is the purest form of cloud waste: full price, zero work. Auto suspend fixes it, but set it too aggressively and you pay in cold caches and resume minimums. Here is where to land.
Databricks DBU cost control: the four levers that matter
A Databricks bill is two bills stacked: DBUs to Databricks and instances to your cloud provider. Controlling it means understanding which compute type you picked, because the DBU rate varies by more than 5x.
Job clusters vs all purpose clusters: a 3.7x cost difference for identical work
The same notebook, the same data, the same runtime, and a bill that differs by a factor of nearly four. The cluster type you attach a scheduled job to is the cheapest big win in Databricks.
Pub/Sub cost explained: what $40 per TiB actually buys you
Pub/Sub prices throughput at a flat rate per TiB with no shards to provision, which is either a bargain or a trap depending on your fan out factor and retention settings. Here is the full meter list.
ClickHouse cost vs a managed warehouse: when self hosting pays off
ClickHouse on your own instances can serve analytical queries for a fraction of per query warehouse pricing, but only above a certain query volume. Here is where the crossover sits and what it costs to cross it.
Elasticsearch vs OpenSearch cost: comparing search cluster bills honestly
Both run the same shaped cluster on the same instances, so the infrastructure bill starts similar. The differences show up in licensing, tiering features, and serverless options. Here is the full comparison.
Compression codec cost trade offs: picking between ZSTD, Snappy, and gzip on price
Codec choice moves three meters at once: storage bytes, scan bytes, and CPU seconds. The cheapest codec on storage is rarely the cheapest overall. Here is how to run the comparison on dollars.
dbt materialization cost: view, table, incremental, and what each one bills
The materialization keyword on a dbt model decides whether you pay at build time, at query time, or both. Getting it wrong on a large model can cost thousands a month in unnecessary rebuilds.
CDC pipeline cost: what change data capture actually charges you for
Change data capture replaces expensive full table reloads with a stream of row level changes, but it introduces its own meters: connector compute, log retention, stream throughput, and merge cost at the sink.
Managed data ingestion cost models: per row, per GB, per connector, per credit
Managed ingestion platforms price the same work four completely different ways, and the cheapest model depends entirely on the shape of your data. Here is how to compare them without guessing.
Analytics sandbox cost: why your non production data stack costs as much as production
Sandbox warehouses, dev clusters, and staging pipelines quietly replicate production capacity for a fraction of the value. Scoping them properly typically reclaims 60 to 80 percent of non production spend.
Data transfer cost between analytics services: the line nobody budgets for
Moving data between a lake, a warehouse, a stream, and a BI tool crosses availability zones, regions, and account boundaries. Each crossing has a price, and at analytics volumes they add up fast.