Query, embedding and prompt caching
Three different caches with the same name. A query cache stores the finished answer, keyed on the question. An embedding cache stores a vector, keyed on the text. A prompt cache stores the model's internal KV state for a shared prefix. They save different things and go stale in different ways.
Overview
Embedding cache: the easy one
Keyed on a hash of the text and the model name. Embedding is deterministic, so the same text always gives the same vector — which makes this cache both trivially correct and permanently valid, until you change models.
Include the model identifier in the key. Vectors from different models are not comparable, and a cache that silently mixes them produces retrieval results that look plausible and are meaningless. The biggest win is at indexing time: re-indexing a corpus after a chunking tweak re-embeds only the chunks that actually changed.
Parameters
Visualisation
—Readout
What to watch
- Each cache is keyed on something different, so each has a different hit rate.
- An exact repeat hits the query cache and skips everything.
- A reworded question misses both exact caches and can still hit the prompt cache.
Query, embedding and prompt caching: A Practical Guide
What are the different caches in a RAG system, and what does each one actually save?
Query cache: the highest saving, and the highest risk
Keyed on the question, storing the final answer. A hit skips retrieval, reranking and generation — often seconds and most of the cost. Real traffic is heavily repetitive, so hit rates can be high.
Two problems. Staleness: the answer was correct for the corpus as it was, so any document update can invalidate it, and there is no cheap way to know which entries. Most systems use a short TTL and flush on re-index. Exact matching: "what is the refund policy" and "how do refunds work" are one question and two keys. Semantic caching — keying on the embedding and accepting a near match — raises the hit rate and introduces the risk of returning the answer to a subtly different question.
Prompt caching: inside the model
Provider-side, and a different mechanism entirely: the model keeps the attention state (the KV cache) for a prefix it has already processed. Send the same long system prompt and the prefill for that portion is skipped.
It is prefix-based, so it only helps if the shared part comes first. Put the system prompt and any fixed instructions at the front and the retrieved chunks and the question after them; putting the variable part first defeats it entirely. Typical savings are large on time-to-first-token and on input cost, and nothing about correctness changes — the model computes the same thing.
Semantic caching, and the risk it introduces
An exact-match query cache misses on any rewording, and real users reword constantly. Semantic caching keys on the question's embedding instead and returns a stored answer when a previous question is close enough.
It raises the hit rate substantially and introduces a failure the exact cache cannot have: returning the answer to a similar but different question. "What is the refund policy for digital goods?" and "What is the refund policy?" are close in embedding space and have different answers. The threshold is the whole design, and it should be set from measured false-hit rate on real traffic, not chosen.
Two safeguards are worth the effort: exclude anything user-specific or permission-scoped from the cache entirely, and log hits so a wrong answer can be traced back to the question that seeded it.
Invalidation, which is where these systems actually break
The embedding cache never goes stale — embedding is deterministic, so the same text and model always give the same vector. Include the model name in the key and it is correct forever.
The query cache is the opposite. It stores a conclusion drawn from a corpus at a moment in time, and nothing in it knows when a document changed. The practical options are all imperfect: a short TTL, which trades hit rate for staleness; a flush on re-index, which is coarse but honest; or tracking which chunks contributed to each cached answer and invalidating precisely, which is correct and rarely worth the bookkeeping.
Prompt caching sidesteps the question because it caches computation rather than conclusions. The model recomputes the same thing it would have anyway, so there is nothing to go stale — which is why it is the safest of the three and the one to reach for first.
Four places a RAG pipeline repeats work
A retrieval pipeline does the same expensive things over and over. Each stage can be cached, and they have very different hit rates and risks.
| Layer | What is cached | Keyed on | Typical saving |
|---|---|---|---|
| Embedding | Query vectors | Normalised query text | One model call |
| Retrieval | Result ids | Query + filters | One index search |
| Reranking | Scores | Query + document id | One cross-encoder pass |
| Generation | Final answers | Query + context | The whole LLM call |
| Prompt prefix | KV cache for a shared prefix | Prefix tokens | Prefill compute |
The last row is different in kind and is the one with the best return in production: if every request begins with the same long system prompt, its KV cache can be computed once and reused, removing that prefill work from every subsequent request.
Exact caching, and its keys
The simplest and safest form. Normalise the query, hash it, look up the result.
import hashlib, json
def cache_key(query, filters, user_groups):
payload = json.dumps({
"q": " ".join(query.lower().split()), # normalise whitespace and case
"f": filters,
"g": sorted(user_groups), # permissions are part of the key
}, sort_keys=True)
return hashlib.sha256(payload.encode()).hexdigest()Two things in that function are the ones that go wrong.
Normalise the query text, or "How do I reset my password?" and "how do i reset my password" are separate entries and the hit rate collapses.
Include the permission context in the key. This is not an optimisation detail — a cache keyed only on query text will serve one user's authorised results to another user who is not authorised to see them. It is the most common way a correctly-built permission filter is defeated, because caching is added later by someone optimising latency.
Semantic caching, and why it is risky
Exact caching only helps on repeated queries. Semantic caching extends it: embed the query, and if it is within some similarity threshold of a cached query, return that entry.
The appeal is a much higher hit rate — "reset password", "how to change my password" and "forgot password" all hit one entry.
The risk is that similar is not the same. Consider:
| Query A | Query B | Cosine | Same answer? |
|---|---|---|---|
| "refund policy" | "refund policy for digital goods" | 0.92 | No |
| "how to cancel" | "how to cancel a subscription" | 0.94 | Maybe |
| "is X supported" | "is X not supported" | 0.96 | No |
The negation row is the dangerous one, and it is a known weakness of embeddings: "is X supported" and "is X not supported" are near-identical vectors and opposite questions.
So semantic caching needs a high threshold (0.97 and above), and it is unsuitable wherever a wrong answer is costly. A safer variant caches only the retrieval stage semantically — returning similar documents for a similar query is far more defensible than returning a previous answer.
Three caches, three hit rates, three ways to serve a stale answer
A RAG pipeline has several independent places to cache, and they differ in what they key on, how often they hit, and how badly they fail when the underlying data changes. This measures all three on the same traffic.
Things to try
- Push exact repeats to 70%. The query cache dominates — and it is the cache that goes stale the moment a document changes.
- Set exact repeats to 0 and reworded to 60%. The exact caches stop helping entirely; only the prompt cache, which keys on the shared prefix, still does.
- Turn prompt caching off. Reworded queries lose their only remaining saving, which is why prompt order matters.
What to remember
Three different caches share the name. The query cache stores a finished answer keyed on the question and saves the most per hit, at the risk of serving a stale one. The embedding cache stores a vector keyed on the text and is permanently valid until the model changes. Prompt caching lives in the provider and only helps when the shared text comes first in the prompt.
Invalidation, which is the hard part
A cache that never expires eventually serves answers from documents that have been superseded — which is a correctness failure, not a performance one.
Four strategies, in increasing order of effort:
Time-based (TTL). Entries expire after an hour, a day, a week. Simple, and it guarantees a staleness window.
Version-based. Include a corpus version or index build id in the cache key. Re-indexing changes the version, so every entry misses and is naturally replaced. Clean and effective.
Event-based. When a document changes, invalidate the cached entries that cited it. Requires recording which documents each entry used — worth doing, and it is the only approach that is both fresh and efficient.
Manual. A flush endpoint for when something has gone wrong. Always have one.
The version-in-the-key approach is the one to reach for by default, because it makes staleness structurally impossible after a re-index rather than relying on a timer.
What is worth caching, and what is not
Cache the embedding layer always. Deterministic, cheap to store, no correctness risk, and query repetition is high in real traffic.
Cache retrieval results with the permission context and corpus version in the key.
Cache the prompt prefix's KV if you control the serving stack. A long shared system prompt or a fixed instruction block is prefilled once instead of per request.
Be cautious caching generated answers. The saving is largest and so is the risk: a stale or mis-keyed answer is served with full confidence and a citation. Use exact keys, a short TTL, and a corpus version.
Do not cache when the answer depends on time ("what is outstanding today"), on per-user data beyond permissions, or on anything with legal or safety weight.
Measuring it
Three numbers tell you whether the cache is earning its place:
Hit rate per layer. Embedding caches often reach 30–60% on real traffic; generation caches much less, because full queries repeat rarely.
Latency at the percentiles. A cache improves the median dramatically and may not move p99, which is what users notice.
Cost saved, in model calls avoided — usually the number that justifies the work.
And one to watch: staleness incidents. If nobody is measuring how often a cached answer was wrong, the cache is untested. A sampled comparison — recompute a small fraction of cached responses and diff them — catches invalidation bugs before users do.
Questions people ask
Where should I start? Embedding caching — safe, cheap, and immediately useful. Then prompt-prefix caching if you control serving.
Is semantic caching safe? Only with a high threshold, and preferably applied to retrieval rather than to final answers. Negation and qualifiers break it.
How do I stop a cache leaking across users? Include the permission context in the key. Nothing else is sufficient.
What TTL? Matched to how often the corpus changes — and prefer a corpus version in the key over a timer.
Does caching hurt quality? Only through staleness and mis-keying, and both are avoidable. The cached computation is identical.
What hit rate should I expect? Highly traffic-dependent. Embeddings often 30–60%; exact answer caches frequently under 10%.
Recap in one screen
- Cache embeddings, retrieval results, rerank scores, answers and the shared prompt prefix — different savings and risks each.
- Normalise the query text before hashing, or the hit rate collapses.
- Include the permission context in every key, or the cache defeats your access control.
- Put the corpus version in the key so re-indexing invalidates naturally.
- Semantic caching raises hit rates and breaks on negation and qualifiers — use a high threshold, and prefer caching retrieval over answers.