LLM Caching
One of the simplest ways to save a lot of money and speed up responses: don't pay twice for the same computation. Three types of caching and when to use each.
Why caching matters
Every LLM call costs money (by tokens) and time. But a large share of calls repeats — the same long system prompt, the same fixed context, or even the same question asked again and again. Caching means: compute once, reuse. The result — cost savings (sometimes tens of percent) and much faster responses.
3 types: Prompt caching (the fixed part of the prompt), Semantic caching (similar questions), and Exact-match (identical questions). Each for a different scenario.
1. Prompt Caching (Context Caching)
Leading LLM providers let you cache the fixed, long part of the prompt — the system prompt, instructions, a context document that recurs on every call. On subsequent calls, that part isn't recomputed and is billed at a fraction of the price.
When it's gold: a chatbot with a long system prompt, RAG with a fixed document, or an agent that sends the same instructions repeatedly. Example with Anthropic:
import anthropic
client = anthropic.Anthropic()
resp = client.messages.create(
model=MODEL, # your provider's current model id
max_tokens=1024,
system=[{
"type": "text",
"text": LONG_SYSTEM_PROMPT, # instructions + fixed context
"cache_control": {"type": "ephemeral"} # mark for caching
}],
messages=[{"role": "user", "content": user_question}],
)
# the first call builds the cache; the following ones use it and are much cheaper
- Put the fixed part at the start (the cache works on a shared prefix).
- The cache is usually short-lived (minutes) — effective for continuous traffic.
- With OpenAI it's usually automatic for long, repeated prompts.
2. Semantic Caching
Here you save the entire call: if a question similar in meaning was already answered — return the cached answer without calling the LLM at all. You use embeddings: compute a vector for the question, search the cache for a question with cosine similarity above a threshold, and if found — return it.
def semantic_cache_get(question, cache, threshold=0.92):
q_vec = embed(question)
for entry in cache: # in production: a vector DB, not a loop
if cosine(q_vec, entry["vec"]) >= threshold:
return entry["answer"] # cache hit — no LLM call
return None
answer = semantic_cache_get(q, cache)
if answer is None:
answer = call_llm(q) # cache miss
cache.append({"vec": embed(q), "answer": answer, "q": q})
- When it's great: FAQ, support, common questions that recur in different phrasings.
- The threshold is critical: too high — misses; too low — inaccurate answers. Calibrate (~0.9+).
- Careful with personalization: don't return a cached answer when the answer depends on a specific user/context.
3. Exact-match Caching
The simplest: the cache key = a hash of the full input (prompt + parameters). If exactly the same input recurs — return the stored output. Fast and cheap, but only catches exact repeats. Good for deterministic tasks (temperature 0) that recur exactly.
import hashlib, json
def key(payload): return hashlib.sha256(json.dumps(payload, sort_keys=True).encode()).hexdigest()
k = key({"model":MODEL,"temp":0,"prompt":prompt})
if k in store: return store[k]
out = call_llm(prompt); store[k] = out; return out
Cache invalidation
A stale cache = wrong answers. Manage it:
- TTL: automatic expiry (hours/days) based on how much the information changes.
- Content version: if you updated the knowledge base/prompt — invalidate the relevant cache (e.g. include a version in the key).
- Don't store sensitive/personal information in a shared cache.
Prompt caching is prefix matching, and that governs everything
The mechanic behind provider prompt caching is simple and it dictates how you must structure every prompt: the cache matches on a shared prefix. Everything from the start of the prompt up to the first difference can be reused; from the first difference onward, nothing can.
One consequence dominates all the others. A single changed character near the top invalidates the entire cache below it. Put a timestamp, a user name, a session id or a randomly ordered list of retrieved documents at the beginning, and you have a cache that never hits while still paying the cost of trying.
So the ordering rule is absolute: most static first, most dynamic last.
- System instructions that never change.
- Tool and schema definitions.
- Few-shot examples.
- Long reference documents that are the same for every user.
- Per-tenant or per-user context.
- Conversation history.
- The current question.
Two practical corollaries. Sort anything you assemble — if retrieved passages arrive in a different order each time, the prompt differs even when the content is identical, so sort them by id before inserting. And do not edit the system prompt casually in a high-traffic system: every change resets the cache for everyone until it warms again.
Caching a prompt used once costs you money
Providers generally charge a premium to write a cache entry and a steep discount to read one. That makes the economics a break-even calculation rather than a free win.
The logic: the first call costs more than an uncached call. Every subsequent call within the cache's lifetime costs much less. You come out ahead once enough reads amortise the one write — often just a couple of hits, sometimes more, depending on the provider's ratio.
Which tells you exactly where to use it:
- Worth caching: a long system prompt on a busy endpoint; a fixed reference document every user's request includes; a tool schema sent on every agent step.
- Not worth caching: anything called a handful of times a day, because the entry expires before it is reused. You pay the write premium repeatedly and read it rarely.
- Not worth caching: short prompts. The saving is proportional to the cached length, and below a provider's minimum it will not cache at all.
Check your provider's write premium, read discount and minimum cacheable length before designing around it — those three numbers determine whether this is a large saving or a rounding error for your traffic.
Cache lifetime meets your traffic pattern
Provider caches are short-lived by design — minutes rather than days, with longer options sometimes available at a higher price. That interacts with your traffic in a way worth thinking about before you attribute a poor hit rate to the feature.
Continuous traffic keeps the entry warm and the hit rate high. This is the case the feature was built for.
Bursty traffic — a few requests an hour, or a batch job that runs twice a day — mostly misses. The entry has expired between uses, so nearly every call pays the write premium and reads nothing back.
Two ways to improve a bursty pattern. Batch the work so calls arrive close together rather than spread out; processing a hundred documents in one run rather than one an hour turns a miss-heavy workload into a hit-heavy one. And where a provider offers a longer-lived cache at extra cost, compare that premium against your actual inter-request gap rather than assuming the default is right.
The risk that makes semantic caching different
Prompt caching cannot give a wrong answer — it is the same computation, cheaper. Semantic caching can, and this is the distinction to hold on to when deciding where to use it.
Returning a stored answer for a "similar enough" question means a similarity threshold is deciding correctness. Consider two questions that sit very close in embedding space:
- "How do I cancel my subscription?" and "How do I cancel my order?" — different processes, different teams, different outcome for the user.
- "Is the pro plan included?" and "Is the pro plan excluded?" — near-identical vectors, opposite meanings. Embeddings handle negation poorly.
- "What is the return window for shoes?" and "…for electronics?" — the same question shape with a different answer.
Raise the threshold and you get fewer hits but safer ones; lower it and you save more money while occasionally answering the wrong question confidently. That is a genuine trade with user harm on one side, not a tuning parameter to optimise.
Two habits make it survivable. Log every cache hit with both questions — the one asked and the one whose answer was served — and read a sample weekly; wrong hits are obvious to a human and invisible to a metric. And never cache semantically where being wrong is expensive: pricing, eligibility, medical, legal, anything about a specific account.
The cache key is a security boundary
The worst cache bug is not staleness — it is one customer receiving an answer computed for another. It happens whenever the key omits something the answer depended on.
An answer is a function of every input, so the key must include every input:
- Tenant or account id, always, for any system serving more than one organisation. This one is non-negotiable and it is the common breach.
- The user's permissions, or a hash of them. Two people in the same company may be entitled to different answers, and the cache does not know that.
- Language and locale, which change the answer and are easy to forget.
- The model and parameters, since a different model produces a different answer to the same prompt.
- A version of the underlying content, so that updating a document invalidates what was derived from it.
A useful rule of thumb: if two requests would legitimately get different answers, they must have different keys. The cheapest safe default for a multi-tenant product is a separate cache namespace per tenant — slightly lower hit rates, and a category of incident that becomes structurally impossible.
Invalidation, and why it is the hard part
The existing advice — TTL and content versions — is right, and the difficulty is in the detail of the second one.
A cached answer derived from three retrieved documents is stale the moment any of them changes. A TTL handles that eventually and badly: too short and the cache does nothing, too long and you serve an answer that contradicts your own documentation for a day.
The better approach is derivation-aware keys: record which source documents an answer used, and when a document changes, invalidate everything derived from it. It costs a small index of answer-to-source relationships and it turns invalidation from a guess into an event.
Where that is too much machinery, the pragmatic compromise is a global version stamp in every key — bump it when you republish the knowledge base or change the prompt, and the whole cache turns over at once. Crude, cheap, and far better than a TTL for content that changes in batches.
One habit worth having regardless: a way to flush the cache immediately, and someone who knows it exists. The first time you publish a correction and the old answer keeps appearing, you will want it.
Which cache to reach for, in order
The three types are not alternatives — they stack, and they should be adopted in a particular order, because the risk and the effort differ enormously.
Start with prompt caching. It is the safest thing on this page: the computation is identical, so it cannot change an answer. It usually requires reordering your prompt and adding one parameter. For most systems with a long system prompt, this is the majority of the available saving and it carries essentially no correctness risk.
Then exact-match caching, for genuinely repeated identical inputs — a batch job that reprocesses the same documents, a public FAQ endpoint, deterministic extraction over a stable corpus. Also safe, provided the key is complete. Normalise the input first — trim whitespace, lowercase where it does not change meaning, sort any lists — or trivially different inputs will miss.
Consider semantic caching last, and only where the questions genuinely repeat in varied wording and being slightly wrong is survivable. Public support content for a product with a stable answer set is the good case. An assistant answering about a customer's own account is not.
A great many teams reach for semantic caching first because it is the interesting one, and take on a correctness risk to capture a saving that prompt caching would have delivered safely.
What to measure, beyond hit rate
Hit rate is the obvious metric and it is not the goal — a cache that returns wrong answers has an excellent hit rate.
- Cost actually saved, against what the same traffic would have cost uncached. This is the number that justifies the work, and it accounts for the write premium that a hit rate ignores.
- Latency at the median and the tail. Caching usually improves the median substantially and does nothing for a miss, so an average hides the shape.
- Stale-hit incidents — how often a cached answer turned out to be wrong or out of date. Rare, high-impact, and only visible if you count it.
- Semantic near-miss rate, from the sampled log of hits described above.
- Cache size and eviction, so an unbounded store does not quietly become your largest infrastructure cost.
Caching and how fast it feels
Cost is the usual reason to cache and latency is the one users notice, so it is worth separating what caching does to each.
Prompt caching cuts the time spent processing the input, which matters most when the prompt is long — a big system prompt plus retrieved documents can dominate the wait before the first word appears. It does nothing for generation time, because the output still has to be produced token by token.
A full response cache — exact or semantic — removes the generation too, which is why a hit feels instant rather than merely quicker. That creates an odd experience worth designing around: some answers arrive immediately and others take several seconds, and the inconsistency reads as unreliability. Streaming the cached answer at a natural pace rather than dumping it at once is a small touch that makes the interface feel coherent.
One more interaction: streaming and caching pull in different directions at the point of writing the cache. You cannot store a response until it is complete, so remember to write the entry after the stream finishes — and handle the case where the user disconnects halfway, which otherwise leaves a truncated answer in the cache to be served to somebody else.
When not to cache
- Creative or varied output. If the point is a fresh draft each time, a cache defeats the feature.
- Personalised answers, unless the key contains the whole personalisation — at which point the hit rate is usually near zero anyway.
- Time-sensitive information — stock, availability, prices, anything with "current" in the question.
- Very low traffic. Below a certain volume the entries expire before reuse and you are paying the write premium for nothing.
- Before you have measured. Caching is an optimisation; applying it to a system whose cost you have not broken down usually optimises the wrong call.
Common mistakes
- Semantic cache with too low a threshold. Returns the answer of an "almost similar" question and hurts accuracy.
- Storing user-dependent answers in a global cache — one customer sees another's answer.
- A cache without a TTL. Answers go stale and become wrong.
- Putting the dynamic part at the start. Breaks prompt caching — the fixed part must be the prefix.
- Forgetting to measure hit rate. Without measurement you don't know if the cache even helps.
Next step
Caching is part of broader cost control. Combine it with model routing and smart pricing.