AI Unit Economics, FinOps & Infrastructure Cost ModelingPlaybook3 min readUpdated September 2026

How Prompt Caching Actually Cuts Your LLM API Bill

Prompt caching is one of the few AI cost levers that's almost entirely free once it's implemented: it costs engineering time to set up, and after that every cache hit is a request you're not paying full price for. The catch is that cache hit rate is fragile and drifts down quietly if nobody's watching it.

Here's how the saving actually works, worked through with a concrete example, and what tends to erode it after launch.

How caching actually reduces the bill

Most model providers charge a lower rate for tokens that were already processed in a recent, matching request, typically the shared prefix of a prompt: system instructions, a knowledge base excerpt, or conversation history repeated across turns. A cache hit means that shared portion is billed at the reduced rate instead of the full rate. The saving scales with how much of your prompt is repeated content versus genuinely new content on each call.

Where cache hit rate actually comes from

Hit rate depends on how much of your prompt structure is stable across requests and how close together in time those requests arrive, since most providers only cache for a limited window. A support chatbot with a long, static system prompt and frequent messages in the same conversation will see a high hit rate. A one-off request with a unique prompt every time will see close to none, no matter how the caching is configured.

A worked example

Say your support bot sends a 2,000 token system prompt plus a short user question on every call, and conversations average six turns. On the first turn, the whole prompt is new and billed at full rate. On the next five turns, the system prompt portion, most of the token count, can be served from cache at the reduced rate, while only the new turn's content is billed in full. Across a full conversation, the blended cost per turn drops substantially once you look past the first message, purely from restructuring the prompt so the stable part comes first and stays byte for byte identical between calls.

The same logic applies to anything else that repeats across calls: a knowledge base excerpt injected into every request, a fixed set of tool definitions, or a style guide the model is told to follow. Anything static that currently gets reassembled and sent fresh on every call is a candidate to move to the front of the prompt and keep identical, purely to make it eligible for the cached rate.

What quietly breaks it in production

Cache hit rate erodes for reasons that rarely show up in a code review:

  • A timestamp, request ID, or other unique value inserted anywhere in the cached portion of the prompt
  • Reordering or lightly editing the system prompt during a routine update, which changes it just enough to miss the cache
  • Conversation gaps longer than the provider's cache window, common in support tools where a customer replies hours later
  • Running the same logical prompt through slightly different formatting in different parts of the codebase

Tracking it as a number finance watches

Cache hit rate should sit next to token spend on whatever dashboard tracks AI cost, not live only in an engineering dashboard nobody outside the team checks. A quiet drop in hit rate looks identical to a rise in traffic on the invoice, and the two need very different responses: one is a cost containment fix, the other is a usage growth story. Without hit rate visible next to spend, finance has no way to tell which one actually happened.

When a low hit rate is the wrong thing to chase

Not every workload can reach a high cache hit rate no matter how well the prompt is structured, and treating a stubbornly low number as a problem to fix can waste engineering time chasing a target the workload was never going to hit. A one-off research query, a document summarization job where every document is different, or a batch classification run over unique records all have little or no repeated content to cache in the first place.

For those workloads, the lever that actually moves cost sits somewhere else: a smaller or cheaper model for the task, a shorter prompt, or batching several similar requests into fewer calls if the use case allows it. Spend the analysis time figuring out which lever fits the workload rather than assuming caching applies everywhere token cost shows up.

The useful question isn't why a hit rate is low in isolation, it's whether the workload has a structurally cacheable shape at all. A support chatbot with a long static system prompt clearly does. A one-off research assistant answering a different question every time from a different context clearly doesn't, and no amount of prompt restructuring will change that.

Executive Capability Standard

What Good Looks Like

Good looks like a stable or improving cache hit rate tracked alongside token spend, with a known owner when it drops.

Building The Capability (5-Stage Skill Ladder)

1. Learn:Understand exactly which part of your prompt structure is meant to be cached and why, before troubleshooting a hit rate problem.
2. Do Manually:Pull cache hit rate from your provider's usage dashboard each week and compare it against the prior week by hand.
3. Delegate:Have the engineer who owns the AI integration report hit rate alongside spend in the same recurring review finance attends.
4. Automate:Add cache hit rate as a tracked field in your usage data pipeline so it shows up automatically next to token spend rather than needing a manual pull.
5. Buy:Consider an LLM observability tool if you're running prompts across multiple features and need hit rate visibility per feature, not just in aggregate.

How to Get Started

Frequently Asked Questions

Why did our cache hit rate drop after we added a new feature?

The most common cause is a prompt structure change: adding a timestamp, a random identifier, or reordered content near the top of the prompt breaks the exact match caching depends on. Check whether anything in the cached portion of the prompt varies between calls that used to be identical.

Does prompt caching affect response quality?

No, caching affects billing and latency for the repeated portion of the prompt, not the content or quality of the model's response. It's a cost and speed optimization, not a tradeoff against accuracy.

How long should we expect a cache to stay warm?

This varies by provider and is usually on the order of minutes, not hours, so caching helps most within an active session and much less for requests spread far apart in time. Check your specific provider's documentation for the current window, since providers do change this.

About the numbers

This guide doesn't quote a sourced benchmark. Figures in it are estimates or general guidance, so check them against your own numbers.

Related Guides