Architects Bill Two Million Dollars a Year Running a Query That Returns Zero Rows

Jul 18, 2026 By Lucas Mendes

In 2025, a senior engineer at a mid-sized fintech company was digging through AWS Cost Explorer when she noticed a line item: $170,000 per month for DynamoDB read capacity units. She traced the spending to a single query pattern executed every 30 seconds by a cron job across 12 microservices. The query returned zero rows. Always had. The job had been running for three years, abandoned by a team that had since been reorganized. The annual cost: just over $2 million for a database read that produced nothing.

Stories like this are not rare. As cloud infrastructure becomes the dominant cost center for technology companies, the gap between what teams think they spend and what they actually burn grows wider. The economics of dead code, orphaned queries, and zombie services represent a quiet tax on every engineering organization. The tax goes uncollected not because it is invisible, but because the incentives to ignore it are stronger than the incentives to fix it.

The Query That Costs $2 Million a Year

The query in question was a simple scan on a DynamoDB table, filtering on a field that no longer existed in any active record. The cron job had been written to pre-warm a cache that was never used. The engineer who wrote it left the company in 2022, and the documentation for the job was a single line in a stale wiki: "Cache warmer for search suggestions." The search suggestion feature had been deprecated six months before the job was deployed.

When the engineer proposed removing the cron job, her team lead hesitated. "What if something downstream depends on it?" The fear is rational. In a microservice architecture of fifty or more services, the call graph is rarely fully known. Observability tools often track requests that return data, but not requests that return nothing. A zero-row query is invisible to most monitoring dashboards. It is not failing, so no alert fires.

The org chart reinforces the status quo. The VP of Engineering approves the capacity budget annually, but the budget is allocated by team, not by feature. The SRE team owns infrastructure cost, but they lack the context to know which queries are essential. No single person owns query lineage. The AWS bill shows aggregate spend, not per-query cost. The $170,000 per month line item was buried under "DynamoDB - ReadCapacityUnits - us-east-1."

By the time the engineer built a simple tagging system to attribute the cost, she found that the same pattern was replicated across 12 microservices. Each service had independently copied the cron job from a shared internal library. No one had ever audited the library's cost footprint. The total waste: roughly $2.1 million per year.

Who Signs the Check for Dead Code

The VP of Engineering signs the capacity budget, but that budget is a lump sum. The SRE team tracks utilization, but they measure aggregate CPU and memory, not query-level efficiency. Finance sees the total cloud bill and compares it to the previous month. A 5% month-over-month increase is flagged; a flat bill is ignored. The $170,000 per month for zero-row queries was flat for three years.

Cost attribution in most organizations is done at the service or team level. A team's monthly spend is the sum of the compute and storage resources they own. But a query that returns zero rows does not appear in any team's cost report unless it is explicitly tagged. The cron job was owned by the "Platform" team, which had been dissolved. The 12 microservices that ran it were owned by 12 different product teams, none of whom knew the job existed.

Finance teams lack the technical context to ask the right questions. They see a line item for DynamoDB reads and assume it corresponds to user-facing traffic. They have no visibility into whether the reads are necessary. The engineer who discovered the waste had to manually cross-reference CloudTrail logs with service ownership maps to trace the cost.

Some organizations have started to address this with internal cost attribution platforms. Stripe, for example, runs quarterly cost attribution reviews where each team presents its infrastructure spend and justifies the largest line items. Netflix uses automated canary analysis to decommission services that have not received traffic in 90 days. Both companies tie infrastructure cost to team P&L, creating a direct incentive to reduce waste.

Why Engineers Do Not Delete

The primary reason engineers do not delete dead code is fear. Fear of breaking something unknown. Fear of being blamed for an outage. Fear of spending two weeks tracing dependencies only to discover that the code is indeed dead, and then getting no credit for the cleanup. The personal risk of deletion outweighs the personal reward.

Observability into query callers is often incomplete. A zero-row query might be polled by a service that has no direct dependency graph. The original author left the company, and the documentation is three years out of date. The engineer considering deletion must either trust the documentation (which is wrong) or trace the code path manually (which is time-consuming). Most choose to leave it running.

Organizational incentives compound the problem. Engineers are rewarded for shipping features, not for reducing cost. A feature that increases revenue by $1 million is celebrated. A cleanup that saves $200,000 per year is a footnote in a quarterly review. The VP of Engineering will not get a bonus for reducing cloud spend by 5%, but they will get a negative review if a deletion causes a production incident.

There is also a cultural dimension. In many engineering teams, code is viewed as an asset. Deleting code feels like destroying value, even when the code has negative value (it costs money to run and maintain). This mindset is slowly shifting as cloud costs grow, but the shift is uneven. As of late 2024, a survey by the Cloud Native Computing Foundation found that only 38% of organizations had a formal process for decommissioning unused services.

The Unseen Economics of Microservice Waste

Zero-row queries are just one type of waste in microservice architectures. Each service tends to duplicate database access patterns. A common example: three different services each maintain their own cache of the same user profile data, refreshing it on independent schedules. The combined read load on the database is three times what is necessary. Shared caches hide this redundancy because each service's cache hit rate looks healthy in isolation.

Orphaned queues are another silent cost. A queue that once processed image uploads continues to consume memory and I/O even after the upload service is deprecated. The queue has no consumers, but the broker still allocates resources to it. In one case, a company found that 30% of its Kafka partitions had no active consumers. The brokers were sized to handle peak throughput that no longer existed.

Idle compute clusters are perhaps the largest category of waste. Many organizations provision clusters for "burst" capacity that is rarely used. A Kubernetes cluster with 100 nodes running at 10% utilization costs the same as one running at 80%. The extra nodes are justified by fear of a traffic spike that never comes. Some estimates put the cost of idle compute in cloud environments at 20-30% of total spend.

The aggregate effect is staggering. A 2025 report from a major cloud consultancy estimated that the average enterprise wastes roughly 25% of its cloud spend on resources that provide no business value. For a company spending $50 million per year on AWS, that is $12.5 million in waste. The zero-row query story is a microcosm of this larger pattern.

How Stripe and Netflix Tackle the Problem

Stripe's approach to cost attribution is instructive. Every quarter, each engineering team presents a breakdown of their infrastructure spend, highlighting the top three cost drivers and explaining why they are necessary. The review is not punitive; it is designed to surface waste that has become invisible. Teams that identify and eliminate waste are recognized in the company-wide engineering newsletter. The result has been a roughly 15% reduction in cloud spend per year, according to a 2024 talk by a Stripe infrastructure engineer.

Netflix takes a more automated approach. Their internal tool, named "Janitor," scans all services and resources for signs of inactivity. If a service has not received any traffic for 90 days, Janitor sends a notification to the owning team. If the team does not respond within two weeks, Janitor decommissions the service automatically. The tool is credited with saving Netflix an estimated $50 million per year in cloud costs, as mentioned in a 2023 engineering blog post.

Both companies tie infrastructure cost to team P&L. When a team's cloud spend is visible in their own budget, they have a direct incentive to optimize. The team that creates the waste owns the cost. This is a stark contrast to the common model where cloud costs are centralized and no single team feels responsible for reducing them.

The key insight is that waste is not a technical problem; it is an incentive problem. Engineers are smart and capable. They will optimize what they are measured on. If they are measured on feature velocity, they will ship features. If they are measured on cost efficiency, they will find waste. The challenge is designing the measurement system.

Three Operational Practices That Stop the Bleeding

The first practice is to instrument every query with caller and purpose tags. A simple middleware that adds a tag like caller:search-suggestions-cron and purpose:cache-warming to every database query makes cost attribution trivial. Tools like AWS X-Ray or OpenTelemetry can propagate these tags automatically. Once every query is tagged, generating a report of the most expensive queries with no callers becomes a simple SQL query.

The second practice is to set expiry dates on all cron jobs and alerts. When a cron job is created, it should have a review date six months in the future. When the review date passes, the job is automatically disabled unless a human explicitly renews it. This forces teams to periodically evaluate whether the job is still necessary. The same principle applies to monitoring alerts: stale alerts that fire on metrics that no longer exist should be deleted.

The third practice is to hold a monthly "zero-row query" review. Once a month, the team gathers to look at the top ten most expensive queries that return zero rows or have no callers. The review is low ceremony: ten minutes, no slides. The goal is to either delete the query or add a reason tag explaining why it must exist. Over time, this habit builds a culture of cost awareness.

Deleting code is best done in pairs: the engineer who understands the code and a peer who understands the downstream dependencies. The pair reviews the dependency graph together and signs off on the deletion. This reduces the personal risk and spreads the knowledge. Budgeting time for cleanup as engineering work, not overhead, is essential. A team that allocates 10% of its sprint capacity to cost optimization will find more waste than a team that treats cleanup as a once-a-year exercise.

The Career Play: Become the Cost Detective

Engineers who find waste get visibility fast. The engineer who discovered the $2 million zero-row query did not just save money; she became the go-to person for cost optimization in her organization. Within six months, she was promoted to staff engineer and given a mandate to build a cost attribution platform. Her career trajectory changed because she read the AWS bill like a debug log.

The skill of reading cloud cost reports is undervalued. Most engineers never look at the AWS Cost Explorer. Those who do often see only aggregate numbers. The cost detective learns to break down spend by API operation, by resource tag, by availability zone. They correlate cost spikes with deployments. They notice when a new feature adds a database query that costs more than the feature's expected revenue.

At Coinbase, an engineer noticed that a single API endpoint was responsible for 40% of the company's database read load. The endpoint was called by a background job that had been written to support a now-deprecated feature. The engineer removed the job and saved the company roughly $4 million per year. He was promoted to senior staff engineer and now leads the platform efficiency team.

A new role is emerging: Cloud Efficiency Engineer. These engineers combine deep infrastructure knowledge with financial acumen. They understand how to instrument systems for cost visibility, how to build cost attribution models, and how to negotiate with cloud providers. The role is still rare, but demand is growing. As cloud costs continue to rise, the engineers who can find and eliminate waste will become increasingly valuable.

The zero-row query story is not an anomaly. It is a symptom of a system designed for growth, not efficiency. The organizations that learn to measure and incentivize cost optimization will have a significant advantage. The engineers who learn to read the cost signals will have a significant career edge. The query that returns zero rows is a signal. The question is whether anyone is listening.

Recommend Posts
Tech

One React Render Architecture Shapes Three UI Team Career Paths

By Sara Park/Jul 18, 2026

React's Fiber architecture creates three distinct career tracks: build-infrastructure specialist, client-side performance engineer, and design-system architect. Each path pays differently and demands different trade-offs.
Tech

One iOS Dev's App Store Review Bypass Took Three Months of Negotiation

By Deepa Iyer/Jul 18, 2026

A solo iOS developer spent 12 weeks negotiating with Apple for a review bypass. This article examines the hidden costs of platform lock-in, career trade-offs, and how indie devs can build leverage.
Tech

One Maintainer's Two-Factor Bypass Was a Flag in an Unread Config File

By Deepa Iyer/Jul 18, 2026

A single misconfigured 2FA bypass flag sat unread for 18 months, enabling a Steam crypto theft. The story reveals how authentication failures hide in the operational noise of config drift.
Tech

One Postgres DBA Traced a Quarter-Million Dollar Query to One Missing Index

By Deepa Iyer/Jul 18, 2026

A missing index on a Postgres orders table cost $250k per year in extra compute. A DBA traced it in weeks. This is the economics of indexing at scale.
Tech

Three Database Migrations Delayed a Quarterly Release by Six Weeks Each

By Lucas Mendes/Jul 18, 2026

Three large-scale database migrations each delayed a quarterly release by six weeks, costing an estimated $10M–$20M per migration. An analysis of the operational failures and business impact.
Tech

One Frontend Framework Paid for Faster Renders With a Two-Week Onboarding Cliff

By Sara Park/Jul 18, 2026

Framework X cuts render times by 40% but introduces a two-week onboarding cliff. Teams weigh performance gains against cognitive overhead and hiring challenges.
Tech

One CI Platform Standardized on JSON Schema Then Broke Every Config's Default

By Sara Park/Jul 18, 2026

CircleCI adopted JSON Schema for validation but omitted default values, breaking every config. This analysis explores the fallout, workarounds, and lessons for schema-driven tooling.
Tech

Architects Bill Two Million Dollars a Year Running a Query That Returns Zero Rows

By Lucas Mendes/Jul 18, 2026

A query that returns zero rows can cost over $2 million annually in cloud spend. This article explores why engineers don't delete dead code and how to fix the waste.
Tech

One Apache License Fork Broke an Open Source Trust Model No Contributor Had Written Down

By Deepa Iyer/Jul 18, 2026

The Redis-to-Valkey fork exposed unwritten rules of open source trust. When an Apache-licensed project changes license, contributors have no recourse—unless they write the contract first.
Tech

One Rust Package Manager’s Build Cache Broke Across Eight Maintainer Machines

By Sara Park/Jul 18, 2026

A corrupted Cargo cache stumped eight maintainers for days. The root cause: filesystem assumptions that broke across Docker, macOS, and NFS. A deep dive into reproducible build challenges.
Tech

One CDN SRE Tracks a Thousand Dollar Spike to a Single Misconfigured Cache Key

By Sara Park/Jul 18, 2026

How a single misconfigured cache key caused a $1,000 CDN spike overnight, and what it reveals about the economics of edge infrastructure in 2026.
Tech

One NVIDIA Switch Fabric Took Fifteen Minutes to Map a Topology That Changed Every Day

By Deepa Iyer/Jul 18, 2026

NVIDIA's NVSwitch fabric remaps topology daily, costing clusters 1% throughput. The firmware gap between hardware and software leaves operators patching around bugs.
Tech

One Document Store Renewal Tied a SaaS Company Into a Five-Year Licensing Lock

By Yusuke Tanaka/Jul 18, 2026

How a SaaS startup's $200k document store migration ballooned to $2.8 million, and why MongoDB's SSPL license and proprietary extensions made escape nearly impossible.
Tech

Platform Fees Fund One iOS Calendar but Block Two Android Widgets

By Deepa Iyer/Jul 17, 2026

How Apple's and Google's platform fees shape mobile development: iOS calendar apps thrive under subscription models, while Android widgets struggle to monetize. A look at the economics behind the code.
Tech

One Firmware Maintainer's Bus Factor Was One Person With One Laptop

By Lucas Mendes/Jul 18, 2026

The story of a single maintainer holding a chip's fate on one laptop. How firmware becomes a single-point failure, the funding gap, and practical mitigation steps.
Tech

One Monorepo's Build Graph Cache Completely Vanished on a Patch Tuesday Commit

By Sara Park/Jul 18, 2026

A Patch Tuesday commit wiped a monorepo's build cache to zero. Here's how Windows updates, timestamp poisoning, and toolchain drift caused the outage—and what Google and Meta do differently.
Tech

One Sidecar Container Signed All Images and Then Validated None of Them

By Deepa Iyer/Jul 18, 2026

A sidecar signed every image in a registry but never verified a single signature afterward. That gap opened a supply-chain attack path that most teams still ignore.
Tech

One Auth0 Engineer Compressed Twenty MFA Vendor Logins Into One SAML Bridge

By Lucas Mendes/Jul 18, 2026

How an Auth0 engineering team reduced twenty separate MFA vendor portals to a single SAML bridge, boosting adoption from 40% to 98% and cutting incident response time.
Tech

One iOS Market Forces Forty Teams to Dual-Write Every Screen

By Sara Park/Jul 18, 2026

An investigation into why forty teams across ten companies maintain parallel iOS and Android codebases, and why cross-platform tools haven't eliminated the dual-write burden.
Tech

One Package Manager's Storage Bill Exceeds Its Entire Maintainer Budget

By Lucas Mendes/Jul 18, 2026

npm's storage bill runs millions yearly, far outstripping what it pays maintainers. The economics of centralized package registries and what can be done.