What is Recall@k?
Of all the documents that should have been retrieved, what fraction actually made the top k? Recall@k = relevant retrieved in top k ÷ total relevant in the corpus. It is the metric that answers the question a RAG pipeline actually cares about — did the evidence reach the model? — and the only one of the four that can be improved simply by looking deeper.
Overview
The definition, and the denominator people forget
Recall@k is the number of relevant documents in the top k divided by the total number of relevant documents that exist. The numerator is easy; the denominator is where the difficulty lives.
Knowing it requires knowing every relevant document in the corpus for every evaluation query. On a corpus of a hundred documents you can label exhaustively. On a million you cannot, so the denominator is estimated — usually by pooling: run several different retrievers, judge the union of everything any of them returned, and treat that pool as the ground truth. It is the standard approach and it is biased, because a document no retriever surfaced is silently counted as irrelevant.
Be honest about this when you report recall. "Recall@5 is 0.83 against a pooled judgement set" is a defensible claim; "recall@5 is 0.83" without qualification implies exhaustive labelling you almost certainly do not have.
Parameters
Visualisation
—Readout
What to watch
- Click results to change which are relevant; the denominator moves too.
- Raising k can only ever increase recall, never decrease it.
- Recall reaches 1.0 only when every relevant document is inside k.
What is Recall@k?: A Practical Guide
What does Recall@k measure, and why is it usually the metric that matters most for a RAG retriever?
Why it is the retrieval metric for RAG
The generator can ignore an irrelevant chunk. It cannot invent a relevant one that was never retrieved. That asymmetry is the whole argument: recall failures are unrecoverable, precision failures are merely expensive.
So the usual configuration is to retrieve generously — a large k, favouring recall — and then let a reranker or the model's own attention handle precision. Retrieve 50, rerank to 5, pass 5. Recall@50 is the number that bounds what the reranker can possibly achieve, and precision@5 is what the generator actually sees.
This is also why hybrid search exists. Dense and sparse retrievers fail on different queries, so taking the union of both raises recall well above either alone — see hybrid search. Fusing two retrievers is a recall strategy first and a ranking strategy second.
The monotonicity that makes it easy to game
Recall@k never decreases as k grows. Retrieve the entire corpus and recall is exactly 1.0. That makes it trivially gameable and means a recall number without its k is not a number at all.
It also means recall must always be read alongside a cost. In classical IR that partner is precision, and the two are combined into F1. In RAG the more meaningful partner is your context budget: recall at the k you can actually afford to put in the prompt. Recall@100 is irrelevant if you pass three chunks.
The useful diagnostic is the gap between recall at a large k and recall at your real k. A large gap means the documents are being found but ranked badly, which is exactly what a cross-encoder reranker fixes. A small gap means better ranking will not help and the problem is upstream — embeddings, chunking, or coverage.
Where recall is the wrong metric
When one document suffices. If a query has a single correct answer, recall@k collapses into Hit Rate@k and the extra machinery buys nothing.
When context is tight. Optimising recall pushes k up, and a long context of mostly-irrelevant chunks measurably degrades generation quality and inflates cost. Past a point, more recall makes answers worse.
When relevance is graded rather than binary. Recall treats a perfect document and a marginally useful one identically. If your judgements have degrees, nDCG uses them and recall throws them away.
The metric that caps everything else
Recall@k asks: of all the documents that would answer this question, what share are in the top k results?
recall@k = (relevant documents retrieved in the top k) / (total relevant documents)
For a RAG system this is the most important retrieval metric, for a structural reason: the model only sees the top k. If the answering chunk is not there, no prompt engineering recovers it. Recall@k is a hard ceiling on the whole system's correctness.
Worked example. A question has 3 relevant chunks in the corpus, and the top 5 results contain 2 of them:
recall@5 = 2/3 = 0.67
When there is exactly one relevant chunk per question — the common case for factoid evaluation sets — recall@k becomes binary per query, and averaging over the set gives the hit rate: the share of questions where the answer was retrieved at all.
Reading the curve, not the number
A single recall figure is much less useful than recall measured at several k.
| k | Recall |
|---|---|
| 1 | 0.42 |
| 3 | 0.61 |
| 5 | 0.71 |
| 10 | 0.86 |
| 20 | 0.94 |
| 50 | 0.96 |
That curve diagnoses the system precisely.
Recall@20 is 0.94 but recall@5 is 0.71. The right chunk is being found and not ranked highly. That is a ranking problem, and a cross-encoder reranker is the fix — retrieve 20, rerank, pass 5.
If recall@50 were 0.60, the document is not being found at all. Reranking cannot help. The fix is upstream: chunking, the embedding model, or adding keyword search.
That distinction — found versus ranked — is the single most useful thing this metric tells you, and it is invisible if you only measure at one k.
The other thing to read from the curve is where it flattens. Between k = 20 and k = 50 recall gains 0.02, so retrieving 50 candidates for reranking buys almost nothing over 20.
Recall against precision
The two pull in opposite directions, and for RAG they are not equally important.
precision@k = (relevant documents in the top k) / k
Increasing k always increases recall (or leaves it unchanged) and usually decreases precision. Retrieve everything and recall is 1.0 with useless precision.
For RAG the asymmetry is real: a missing chunk is fatal, an extra chunk is merely wasteful. So the retrieval stage should favour recall, and precision is recovered by reranking and by passing only the top few to the model.
That said, precision is not free. Irrelevant context measurably degrades answers — the model may use it, and it dilutes attention. So the shape that works is: high recall at a generous k in stage one, then a reranker, then a small precise set into the prompt.
| Stage | Optimise for | Typical k |
|---|---|---|
| Retrieval | Recall | 20–100 |
| Reranking | Precision | Reduces to 3–10 |
| Generation | Uses what it is given | 3–10 |
Of what exists, how much you found
Recall@k is the ceiling on everything downstream: a RAG system cannot answer from a document it never retrieved. This measures where that ceiling sits, what raising k costs, and why recall is the metric to optimise first and report last.
Things to try
- Drag k from 1 to 10 and watch recall climb and never fall. That monotonicity is why a recall figure without its k means nothing.
- Mark one more result relevant. Recall drops even though nothing about the ranking changed — the denominator grew. This is what makes an incomplete judgement set flatter your system.
- Set k to 3 and compare recall with precision. They move in opposite directions as k changes, which is the trade the whole field is organised around.
What to remember
Recall@k is the fraction of all relevant documents that reached the top k. It is the metric that bounds a RAG pipeline, because a document that was never retrieved cannot be used, while an irrelevant one can be ignored. It rises monotonically with k, so it is meaningless without its k and must be read against a cost — usually your context budget. Its denominator requires knowing every relevant document, which on a real corpus means pooled judgements and an honest caveat.
All four of these read the same object: an ordered list of retrieved results, with each one labelled relevant or not by a human or a strong model. They differ only in what they choose to notice about it — which is why quoting one without saying which is close to meaningless, and why the visualisation above shows all four at once.
The labels are the expensive part. A relevance judgement per query-document pair is human work, and an evaluation set of fifty queries with judged results is worth more than any amount of metric sophistication on top of unjudged data. Build the set first.
Measuring it
You need questions with known relevant chunks. That evaluation set is the work, and everything else follows from it.
def recall_at_k(retrieved_ids, relevant_ids, k):
top_k = set(retrieved_ids[:k])
return len(top_k & set(relevant_ids)) / len(relevant_ids)
def mean_recall(eval_set, retriever, k):
return sum(
recall_at_k(retriever(q["question"], k=k), q["relevant_ids"], k)
for q in eval_set
) / len(eval_set)
for k in (1, 3, 5, 10, 20, 50):
print(k, round(mean_recall(eval_set, retrieve, k), 3))Three routes to building the set:
Hand-written by someone who knows the domain. Highest quality; 50 questions is enough to catch large regressions.
Generated by asking a model to write a question for each chunk. Fast, and it flatters retrieval — the question is phrased like the document, so the vocabulary gap that causes real failures is absent by construction. Expect real-world recall to be lower.
Mined from logs — real queries paired with the documents users clicked, cited or rated helpful. Most realistic, and it needs a system already in production.
Include some unanswerable questions, so you can also measure whether the system declines rather than inventing.
What moves the number
| Change | Effect on recall |
|---|---|
| Add BM25 alongside dense search | Usually the largest single gain |
| Better chunking (structure-aware, overlap) | Substantial |
| A stronger or domain-tuned embedding model | Substantial |
| Query rewriting or multi-query | Helps on vague queries |
| Increase k | Always helps, at a cost |
| Add a reranker | None — it improves ranking, not recall |
That last row is worth stating plainly because it is a common confusion: a reranker cannot increase recall@k for the same k. It reorders the candidates it was given. It improves recall@5 only in the sense that it moves relevant documents from rank 15 into the top 5 — which is exactly why the diagnosis "good recall@20, poor recall@5" points at reranking.
Caveats
It depends entirely on the labels. If your evaluation set marks one chunk as relevant when three are, recall is understated. Marking all acceptable chunks is the fix, and it takes judgement.
It ignores order within the top k. A relevant document at rank 1 and at rank 10 count the same. Use MRR or nDCG when position matters — and it does, because models attend more to earlier context.
It says nothing about the answer. Perfect recall with a poor prompt still produces bad answers. Retrieval metrics and generation metrics are both needed.
Generated evaluation sets overstate it. Worth repeating: the gap between generated-question recall and real-query recall is often large, and it is the whole problem query rewriting exists to solve.
Questions people ask
What is a good recall@5? Corpus-dependent. Compare against your own baseline and track it over time rather than aiming at an absolute number.
Should I measure recall or hit rate? They coincide when there is one relevant chunk per question. Use recall when questions have several.
What k should I optimise for? The k you actually pass to the model. That is the ceiling that matters.
How many evaluation questions? 50 to catch regressions; 200 for stable comparisons between similar configurations.
Does a reranker improve recall? Not at fixed k in the retrieval stage. It improves what reaches the model, which is why the two-stage pipeline exists.
Why is production worse than my evaluation? Almost always because real queries are phrased differently from your generated ones.
Recap in one screen
- Recall@k is the share of relevant documents present in the top k, and it caps the whole system.
- Measure it at several k: a gap between recall@20 and recall@5 means a ranking problem, low recall@50 means a finding problem.
- For RAG, favour recall in retrieval and recover precision with a reranker.
- Build 50–200 questions with known relevant chunks, and include unanswerable ones.
- Generated evaluation questions overstate recall, because they share the document's vocabulary.