The $11k cloud bill was mostly the tools you added to watch the cloud bill
DEV Community

The $11k cloud bill was mostly the tools you added to watch the cloud bill

A startup with only six engineers calculated the cost of their monthly cloud bill to be $11,847. The majority of the data generated was not from the application itself, but from the monitoring tools that were observing the application. ## The recursive joke nobody's laughing at We often see the same pattern in the industry. You create a small app. Next, you attach the equipment needed to monitor the app. And eventually, the monitoring equipment becomes more expensive than the app itself. An example of this was documented in a Medium post on August 13, 2026. One team reduced their invoice from $11,847/mo to $3,812/mo by switching to two VMs. They came from deleting the overhead that was policing four services that could run on two boxes. That's the basic idea. First, you pay to operate the device, and then you pay an additional amount to observe its operation. ๐Ÿ™ƒ The numbers are not subtle This is not just one disgruntled engineer. There's a pattern here that is supported by evidence. Software engineer Devrim Ozcay wrote up a January 13, 2026 recap of a 6-microservice Spring Boot system. His AWS bill was $1,850/mo. Three months in, his Datadog bill reached $3,200/mo and peaked at a $12,000 invoice due to container auto-scaling. Ozcay was straightforward about it: Our bill for AWS infrastructure was $1,850. Our bill for Datadog was $3,200. Let that sink in for a moment. We were literally spending more money to monitor our servers than to actually run them. For example, an Aug 12, 2026, engineering report on a nine-microservice and Lambda stack found AWS compute at $18,432 and Datadog at $24,773 because of high-cardinality custom metrics. Also, on November 13, 2025, an a developer forum thread mentioned a firm that invested $52,000 in AWS costs, but paid $97,000 for observability, with $47k going to Datadog, $38k going to Splunk, and $12k going to Sentry. The reason behind it is in how it is priced. Datadog is priced per host and per custom metric. The tab autoscales with you. One SRE in that thread summed up the conversation with leadership: I was asked by leadership why observability was so expensive. I explained that it was because Datadog has a per host cost, and we implement autoscaling. They stared at me as if I had just said something in Martian. ## The mesh tax nobody budgeted for Having observability is important, but you've likely over-orchestrated. On January 24, 2026, performance tests on service mesh sidecars were shared by Kubernetes performance expert Nawaz Dhandala. Your average sidecar proxy for service mesh consumes 50-100MB of RAM, and 10-50m CPU per pod, additionally it adds 2-5ms of latency per network hop. In an article on August 13, 2026, Cloud architect Khimananda Oli expanded on the disadvantages of sidecar containers, stating that they can consume as much as 200MB RAM each and introduce up to 15ms latency. Do the multiplication. A request crossing five microservices can pile on 75ms of latency and burn a full gigabyte of RAM just to route traffic. This is the painful aspect. You break down the application into smaller parts for quick implementation, but then you face performance and latency issues as you integrate and reassemble those parts over a network. ## People are quietly walking it back The industry is taking this seriously. An estimated 42% of organizations that have adopted microservices has consolidated at least some services back into larger deployable units. Yes, you're right. Actually, it wasn't about ideology at all. The main issues were related to the complexity of debugging and the high costs of network latency - which were exactly the problems the new tooling was meant to address in the first place. Here's how I see it: โ†’ The bill isn't a compute problem, it's an architecture problem wearing a compute costume. โ†’ Per-host, per-metric pricing punishes the elasticity you were sold on. โ†’ Four services on Kubernetes with a mesh is often two VMs cosplaying as "scale." โ†’ You can't monitor your way out of a design you didn't need. This doesn't imply that observability is negative. It just indicates that the ratio is off balance. If the observers are more expensive than the entities being observed, the solution is typically not cheaper observers, but rather reducing the number of entities being observed. First, reduce the surface area, and then, you can proceed to instrument the remaining part. Here's what I'd like to ask you: What's the highest observability-to-compute ratio you've ever encountered, and was there an audible gasp when you mentioned it? Top comments (0)

Read on DEV Community ↗ ← Back to News

Comments

No comments yet. Start the discussion.