Where Observability Budgets Actually Leak, and How to Contain Them
Observability tools like Datadog price on volume: how much you log, how many metrics you emit, how many traces you capture. Engineering teams tend to treat "more visibility" as an unambiguous good, which is exactly how the bill grows faster than the infrastructure it's monitoring.
The fix isn't turning observability off, it's finding the specific places volume is being generated without adding real visibility, and containing those without touching the signal engineers actually rely on.
Debug logging that never got turned back down
The single most common leak is debug-level logging turned on during an incident and never turned back down afterward. It's invisible in the product and invisible in a code review, since the log statements themselves were already there; only the verbosity setting changed. A quarterly audit of log level settings by service, compared against what's actually needed in steady state, catches this reliably.
Set an expiration on any verbosity change made during an incident, not just a mental note to revert it later. A ticket or calendar reminder tied to the specific change, due a set number of days after the incident closes, survives the on-call engineer moving to a different project far better than an intention nobody wrote down.
High cardinality metrics tags
A metric tagged with something like a user ID or request ID multiplies the number of unique time series a monitoring tool has to store, often driving cost far more than the number of metrics themselves. This tends to happen when a well-meaning engineer adds a tag for debugging one specific incident and it ships to every service using that metric going forward. Reviewing tag cardinality periodically, not just metric count, is where the real spend usually is.
Tracing sample rates left at development defaults
Default tracing configuration is often set to capture a much higher percentage of requests than production actually needs for useful visibility, especially once traffic volume is well past what the default was tuned for. Lowering the sample rate for high volume, low risk endpoints while keeping full tracing on critical paths preserves the visibility that matters while cutting the volume that doesn't. Adaptive sampling, where the tool automatically captures a higher share of slow or failed requests and a lower share of ordinary fast ones, usually gets you most of this benefit without hand-tuning a rate per endpoint.
Retention windows nobody revisited
Log and trace retention defaults to a period chosen once, early on, and rarely revisited as volume and cost both grow. Most incident investigation happens within days of an event, not months, so a long retention window on high volume data is often paying to store detail nobody will ever look at again. Shortening retention on the highest volume, lowest value data categories is usually the single largest lever available without any loss of day to day visibility.
The containment move that works without cutting visibility
Route what you keep by value, not by volume: full detail and long retention on the handful of services where an outage is expensive, reduced detail and shorter retention everywhere else. This preserves the debugging depth engineers actually need on the systems that matter most, while cutting cost on the long tail of lower stakes services that were generating the bulk of the volume without a matching amount of usefulness.
Make this an explicit, written tiering rather than an implicit habit, so a new service added six months from now gets classified deliberately instead of defaulting to whatever configuration a template happened to ship with. The list of which services sit in which tier is a short document, and it's worth keeping current as the product changes.
Audit these settings each quarter:
- Check for debug level logging that was turned on during an incident and never turned back down.
- Look for metric tags such as user ID or request ID that multiply unique time series, and remove the ones added for a single incident.
- Compare tracing sample rates against what production traffic actually needs, and lower them for high volume, low risk endpoints.
- Shorten log and trace retention on high volume data that nobody reviews once the first few days after an event have passed.
- Keep full detail and long retention only on the few services where an outage is expensive.
Why per-host and per-GB pricing need different playbooks
Observability vendors price on different units, and the containment tactic that works for one doesn't help at all against the other. A tool billed per host or per container mostly responds to infrastructure footprint: consolidating services onto fewer, larger instances or cleaning up abandoned monitoring agents on decommissioned hosts moves that bill more than trimming log volume ever will. A tool billed per gigabyte of ingested data responds to the volume levers already covered here, log level, cardinality, sample rate and retention, and largely ignores how many hosts you're running.
Check the actual contract before assuming which lever applies. Teams sometimes spend real effort tuning log verbosity against a bill that's actually driven by host count, or chase down zombie monitoring agents on a bill that's actually driven by ingested volume, and neither effort moves the number they're aiming at.
If you're on a per-host plan, an audit of monitoring agents installed on infrastructure that's since been decommissioned or downsized is usually the fastest win, since agents rarely get uninstalled as part of a standard decommissioning checklist, and each one left running quietly keeps billing as if the host it's watching still matters.
What Good Looks Like
Good looks like observability spend growing roughly in line with traffic and headcount, with a clear owner for the configuration driving it.
Building The Capability (5-Stage Skill Ladder)
How to Get Started
Frequently Asked Questions
Will cutting log volume make incidents harder to debug?
Cutting volume indiscriminately would, but targeting debug-level logging left on by accident, high cardinality tags, and retention on data nobody reviews doesn't touch the signal engineers rely on during an incident. The goal is removing volume that isn't adding visibility, not removing visibility itself.
Should finance or engineering own observability cost?
Engineering should own the technical changes, since only it knows what debugging needs, while finance owns visibility into the trend. Finance should ask specific questions when the bill grows faster than traffic or headcount, rather than treating the tool as fixed overhead.
How often should we audit observability configuration?
A quarterly review of log levels, tag cardinality, sample rates and retention windows catches most of the drift before it compounds into a large number. Waiting for an annual budget review to notice the trend usually means catching it well after several quarters of avoidable spend, by which point the fix is the same but the sunk cost is larger.
About the numbers
This guide doesn't quote a sourced benchmark. Figures in it are estimates or general guidance, so check them against your own numbers.
Related Guides
Cutting Your CI Bill Without Slowing Down Deploys
Where CI spend actually accumulates across compute minutes, artifact storage, and runner sizing, and a monthly review that catches drift before it grows.
What a Fractional CFO Costs per Month and How to Budget
How fractional CFOs price their work, what pushes a monthly retainer up or down, and how to test a proposal against your G&A budget before you sign.
Calculating a Real Cost-Per-Transaction Number
Why total infrastructure spend hides whether growth is healthy, and how to build a cost-per-transaction number that survives a shifting mix of usage.
What Evaluating Your AI Agent Actually Costs to Run
See where AI agent evaluation cost comes from: judge-model calls, human review and test set upkeep, with a worked run example and ways to keep spend in check.
Build Your Own AI Inference Cost Model in Three Tabs
How to structure a spreadsheet that turns token usage into a real cost per customer, so you can see GPU and API spend before the invoice arrives.
Open Source vs Proprietary LLMs: What Each Choice Costs Your Margin
A CFO's side-by-side look at what open source and proprietary language models actually cost once you include hosting, tuning and engineering time.