What is Precision@k?
Of the k documents you returned, what fraction were actually useful? Precision@k = relevant in top k ÷ k. Where recall asks whether the evidence arrived, precision asks how much rubbish arrived with it — and in a RAG system that rubbish costs tokens, latency, and measurable answer quality.
Overview
The definition, and the denominator that never moves
Precision@k is the number of relevant documents in the top k divided by k. Not by the number of relevant documents in the corpus, and not by the number retrieved — by k, always.
That fixed denominator is what makes precision cheap to measure. You only need judgements for the k results you actually returned, not for the whole corpus. An evaluation set for precision can be built by judging a few hundred query-document pairs; one for recall needs exhaustive or pooled labelling. If you can only afford one, precision is the affordable one.
The consequence of the fixed denominator: if fewer than k documents are relevant in total, precision@k cannot reach 1.0. With two relevant documents and k=5, the ceiling is 0.4. A low precision score is sometimes a property of the query, not a failure of the retriever.
Parameters
Visualisation
—Readout
What to watch
- The denominator is always k, whatever the corpus contains.
- Precision usually falls as k rises — the best results are at the top.
- Click results to see precision and recall move in opposite directions.
What is Precision@k?: A Practical Guide
What does Precision@k measure, and why does it matter more in RAG than people expect?
Why noise is not free in a RAG pipeline
The old intuition — "the model can just ignore irrelevant chunks" — is not quite true, and the ways it fails are worth naming.
Cost and latency. Every retrieved chunk is input tokens on every request. Halving the number of chunks halves that bill and shortens time to first token.
Distraction. Irrelevant context measurably reduces answer quality. A plausible-but-wrong passage is worse than no passage — the model has been handed a reason to be confidently incorrect, and it is grounded in a real retrieved document, which makes the error harder to spot.
Position effects. Attention over a long context is uneven, with the middle attended least. Padding the prompt with irrelevant chunks pushes good evidence into that dead zone.
So precision is not merely an efficiency metric here. Past a point, improving precision improves the answers.
The trade with recall, and where each belongs
Precision and recall pull against each other as k moves. Raising k can only help recall and usually hurts precision, because the highest-scoring results were already at the top and what follows is progressively worse.
The standard combination is F1, the harmonic mean, which punishes an imbalance — a system with precision 1.0 and recall 0.1 scores 0.18, not 0.55. The harmonic mean is used precisely because it refuses to let one strong number hide a weak one.
In a modern RAG stack the two are usually optimised at different stages. The retriever runs at high k for recall; the reranker cuts that to a handful for precision. Measuring recall@50 and precision@5 in the same evaluation tells you which of the two stages is failing, and they are fixed by completely different work.
What precision cannot see
Order within k. Relevant documents at ranks 1 and 2 score the same as relevant documents at ranks 4 and 5. For a generator, that difference is real; precision is blind to it and MRR or nDCG is not.
What was missed. A retriever that returns three perfect documents and misses seven more scores precision 1.0. The metric is silent about the corpus it did not touch, which is exactly why it is never reported alone.
Degrees of usefulness. Binary relevance forces a yes/no on documents that are genuinely partial. If your judgements are graded, nDCG uses that information and precision discards it.
What share of what you returned is useful
precision@k = (relevant documents in the top k) / k
If 3 of the top 5 results are relevant, precision@5 = 0.6. The denominator is always k, whether or not that many relevant documents exist.
That last detail matters and is a known limitation: if only 2 relevant documents exist in the whole corpus, precision@5 cannot exceed 0.4 however perfect the retrieval. For that reason precision@k is best read alongside recall, not alone.
Worked through. Top 5 results for "parental leave entitlement", with relevance marked:
| Rank | Document | Relevant? |
|---|---|---|
| 1 | Parental leave entitlement section | Yes |
| 2 | Parental leave application form | Yes |
| 3 | Annual leave policy | No |
| 4 | Parental leave — return to work | Yes |
| 5 | Expenses policy | No |
precision@5 = 3/5 = 0.6. precision@3 = 2/3 = 0.67. precision@1 = 1.0.
Note that precision@k does not care where the relevant results sit. Ranks 1, 2, 4 and ranks 3, 4, 5 give the same precision@5, which is why rank-sensitive metrics exist.
Why it matters less than recall for RAG — but still matters
For a retrieval system feeding a language model, the two errors are not symmetric. A missing relevant chunk means the answer cannot be produced. An extra irrelevant chunk costs context budget and some attention.
So the retrieval stage should be tuned for recall, and precision is recovered downstream by reranking and by passing only a few chunks to the model.
That does not make precision irrelevant, for three reasons:
Context budget. Passing 20 chunks when 3 are relevant wastes tokens that cost money and latency.
Irrelevant context degrades answers. Models do use what they are given, and contradictory or off-topic passages measurably reduce correctness.
Attention dilution. More context means each relevant chunk gets proportionally less attention, and material in the middle of a long context is attended to less reliably.
So the target is high recall in stage one and high precision in what finally reaches the prompt — which is precisely what a reranker delivers.
Where precision is the primary metric
User-facing search results. Someone scanning a result list judges the system by the first few entries. Precision@1, @3 and @5 are what they experience; recall@50 is invisible to them.
Recommendations. A carousel of five items is judged on whether those five are good.
Alerting and triage. Every false positive costs human review time, so precision is the operational constraint.
Anything with a tight context budget — a small model, a long prompt, or a per-request cost limit.
In those cases the trade runs the other way: favour precision, accept lower recall, and give the user a way to ask for more.
Of what you showed, how much was right
Precision@k is the simplest retrieval metric and the easiest to misread. This computes it on a small result list, shows the three things it deliberately ignores, and puts it beside the metrics that cover for each one.
Things to try
- Drag k up from 1. Precision generally falls while recall rises — that opposition is the whole reason both numbers are quoted together.
- Set k to 10 with four relevant documents. Precision cannot exceed 0.4 however good the ranking is; the ceiling is set by the query, not the retriever.
- Mark the two lowest results relevant and watch precision@5 stay put. Improving what sits below k does nothing for it — which is why k must match what you actually send to the model.
What to remember
Precision@k is the fraction of the returned k that was relevant, with k always as the denominator. It is the cheap metric to label, because it needs judgements only for what you returned. In a RAG pipeline it is not just about efficiency: irrelevant context costs tokens, pushes good evidence into the least-attended part of the prompt, and gives the model plausible material to be wrong with. Optimise recall at the retriever and precision at the reranker, and report both with their k.
All four of these read the same object: an ordered list of retrieved results, with each one labelled relevant or not by a human or a strong model. They differ only in what they choose to notice about it — which is why quoting one without saying which is close to meaningless, and why the visualisation above shows all four at once.
The labels are the expensive part. A relevance judgement per query-document pair is human work, and an evaluation set of fifty queries with judged results is worth more than any amount of metric sophistication on top of unjudged data. Build the set first.
Combining the two
Neither metric alone describes a retrieval system, and there are standard ways to combine them.
F1@k is the harmonic mean of precision@k and recall@k. One number, and it punishes imbalance: precision 1.0 with recall 0.0 gives an F1 of 0, which is the correct verdict.
Precision-recall curve. Sweep k and plot the pair. The shape shows the whole trade-off rather than one point on it.
Average precision (AP) averages the precision at each rank where a relevant document appears — so it rewards putting relevant documents early, unlike plain precision@k. Averaged across queries it gives MAP, a standard information-retrieval measure.
nDCG@k handles graded relevance and discounts by position logarithmically. The most complete measure, and it needs graded judgements rather than binary ones.
| Metric | Sensitive to rank order? | Needs graded labels? |
|---|---|---|
| Precision@k | No | No |
| Recall@k | No | No |
| MRR | Yes — first relevant only | No |
| MAP | Yes | No |
| nDCG@k | Yes | Yes |
For most RAG work, recall@k plus MRR is sufficient. Add nDCG if you have graded judgements and are comparing configurations closely.
Measuring it
def precision_at_k(retrieved, relevant, k):
top_k = retrieved[:k]
return sum(1 for d in top_k if d in relevant) / k
def f1_at_k(retrieved, relevant, k):
p = precision_at_k(retrieved, relevant, k)
r = len(set(retrieved[:k]) & set(relevant)) / len(relevant)
return 2 * p * r / (p + r) if (p + r) else 0.0Two practical notes on interpretation.
Precision@k is bounded by the number of relevant documents. With 2 relevant documents, precision@10 cannot exceed 0.2. Compare precision at a k close to the number of relevant items, or use MAP.
Judgement quality dominates. A document marked irrelevant that a user would have found useful lowers measured precision without lowering real quality. Binary judgements are cheap and blunt; graded ones are expensive and better.
What moves precision
| Change | Effect |
|---|---|
| Add a cross-encoder reranker | Largest gain — this is what rerankers do |
| Reduce k | Mechanically increases precision, lowers recall |
| Better chunking | Fewer fragmentary near-misses |
| Deduplication or MMR | Removes redundant results occupying slots |
| Metadata filtering | Removes structurally irrelevant candidates |
| Hybrid search | Helps both, mostly recall |
Reranking is the direct instrument. A cross-encoder reads the query and each candidate together and scores them far more accurately than a comparison of precomputed vectors, which is exactly the operation that raises precision within a fixed candidate set.
Questions people ask
Precision or recall for RAG? Recall in the retrieval stage, precision in what reaches the prompt. The two-stage pipeline delivers both.
Why is my precision@10 low even though results look good? Probably because fewer than 10 relevant documents exist. Check the denominator problem.
Should I use F1? Useful as a single summary. Report precision and recall separately as well, since the trade-off is the interesting part.
How many chunks should I pass to the model? 3–10. Measure answer correctness against the number passed — it usually peaks and then declines.
Does high precision mean good answers? Necessary and not sufficient. Generation can still fail with perfect retrieval.
What is MAP? Mean average precision — average precision at each relevant rank, averaged over queries. Rank-sensitive, and a better single number than precision@k.
Recap in one screen
- Precision@k is the share of the top k that is relevant, with k always as the denominator.
- For RAG, a missing chunk is fatal and an extra chunk is wasteful — so favour recall first.
- Irrelevant context still costs money, latency and answer quality, so precision matters in what reaches the prompt.
- Precision@k ignores order and is capped by the number of relevant documents; MAP and nDCG fix both.
- Reranking is the direct way to raise precision within a fixed candidate set.