What is Hit Rate@k?

The bluntest of the ranking metrics, and often the most useful one to start with. Hit Rate@k is 1 if at least one relevant document appears in the top k, and 0 otherwise. It does not care how many relevant documents there were, or where in the top k they landed. Averaged over a query set, it answers a single question: how often does retrieval put something useful in front of the generator?

Overview

The definition, and the averaging that hides in it

For a single query, Hit Rate@k is binary: 1 if any of the top k results is relevant, 0 if none is. There is no partial credit. A query whose top 3 contains five relevant documents and a query whose top 3 contains exactly one both score 1.

What people usually mean by "our hit rate is 0.82" is the mean of that binary value across an evaluation set: 82% of queries had something useful in the top k. That framing is worth saying out loud, because it makes the metric's shape obvious — it is a proportion of queries, not a proportion of documents.

You will also see it called recall@k with a single relevant document, and in recommender literature simply hit rate. When every query has exactly one correct answer, Hit Rate@k and Recall@k are the same number, which is why the two get confused.

Parameters

Visualisation

—

Readout

What to watch

  • Click any result to flip whether it is relevant.
  • Hit Rate jumps to 1 the moment one relevant item enters the top k.
  • Move k below the first relevant rank and it collapses to 0.

What is Hit Rate@k?: A Practical Guide

What does Hit Rate@k measure, and when is it the right metric for a retriever?

Why it is the right first metric for RAG

A RAG generator does not need every relevant document. It needs enough grounding to answer, and for most factual questions one good chunk is enough. If the answer is in the context, the model can use it; if it is not, no amount of prompt engineering will recover it.

That makes Hit Rate@k a measure of the ceiling on your whole pipeline. A hit rate of 0.7 at your context budget means 30% of queries are unanswerable no matter how good the generator is, and every hour spent tuning prompts against those queries is wasted. It is the number to establish before anything else, because it tells you whether your problem is retrieval or generation.

It is also cheap to label. Deciding "is there anything useful here?" is much faster for a human annotator than grading every document on a five-point scale, so a hit-rate evaluation set can be built in an afternoon.

What it deliberately ignores

Position. A relevant document at rank 1 and at rank k score identically. That matters more than it sounds: models attend unevenly across a long context, and evidence buried at the bottom of ten chunks is measurably less likely to be used. Hit rate will not show you that; MRR will.

Quantity. One relevant document scores the same as five. For a question needing several sources — "compare our refund policy across regions" — hit rate can be 1.0 while the answer is hopelessly incomplete. Recall@k is the metric that notices.

Noise. Nine irrelevant results alongside one good one still scores 1. Irrelevant context measurably degrades generation and inflates cost, so a high hit rate with low precision is a real failure mode that this metric reports as success.

Choosing k, and reading the curve

k should be the number of chunks you actually put in the prompt. Evaluating at k=10 when you pass 3 to the model measures a system you are not running. This is the single most common mistake with @k metrics.

Plotting hit rate against k is more informative than any single value. If it climbs steeply from k=1 to k=5 and then flattens, your retriever finds the right documents but ranks them poorly — a reranker will help a lot. If it is flat and low from the start, the documents are not being retrieved at all, and the problem is your embeddings, your chunking, or the fact that the answer is not in the corpus.

That diagnostic split is the most valuable thing hit rate gives you, and it costs one evaluation run.

Did we find anything useful at all?

Hit rate@k is the simplest retrieval metric: for what share of queries does at least one relevant document appear in the top k?

hit rate@k = (queries with ≥1 relevant result in the top k) / (total queries)

It is binary per query. Ten relevant documents in the top k counts the same as one, and position within the top k is ignored entirely.

Worked example over five queries at k = 5:

QueryRelevant found in top 5?
1Yes, at rank 2
2Yes, at rank 1
3No
4Yes, at rank 5
5Yes, at ranks 1 and 3

hit rate@5 = 4/5 = 0.80

Note that when each question has exactly one relevant document — the common case for factoid evaluation sets — hit rate@k and recall@k are the same number. The distinction only appears when questions have several relevant documents.

Why the crudeness is sometimes what you want

Hit rate throws away information, and that is occasionally the point.

It is the right framing for question answering. If a single chunk answers the question, finding it once is success. Finding it twice adds nothing, and MRR's penalty for it appearing at rank 3 rather than rank 1 may not reflect anything a user experiences.

It is easy to communicate. "In 87% of cases we retrieve at least one document that answers the question" is a sentence a stakeholder understands immediately. "MRR is 0.64" is not.

It is a clean ceiling. Since the model can only answer from what it was given, hit rate@k where k is the number of chunks passed is an upper bound on the system's answerable share.

MetricCountsPosition mattersSeveral relevant docs
Hit rate@kAny hitNoIgnored
Recall@kShare foundNoCounted
Precision@kShare usefulNoCounted
MRRFirst hitYesIgnored
nDCG@kAll, gradedYesCounted

Reading it across k

As with recall, a single figure is far less useful than the curve.

kHit rate
10.51
30.72
50.81
100.90
200.95

Two readings from that table.

Hit rate@20 of 0.95 with hit rate@5 of 0.81 means 14% of queries have their answer between ranks 6 and 20 — found, but not passed to the model at k = 5. That is a reranking opportunity, and it is the most common actionable finding.

The remaining 5% at k = 20 are outright retrieval failures. Look at those queries individually: they usually cluster around a cause — a vocabulary the embedding model does not cover, a document type that chunked badly, a question the corpus genuinely cannot answer.

That second habit matters more than the aggregate. Twenty failing queries read individually will tell you more about what to fix than any amount of averaging.

The averaging that hides inside one number

Hit rate is the crudest retrieval metric and often the right first one. The section above names the averaging that hides in it -- this measures what that averaging conceals, and what hit rate stops being able to tell you as k grows.

example_01.pyNumPy
Output

Things to try

  1. Set k to 1. Hit rate is 1 only if the very first result is relevant — this is the strictest possible reading, and the closest to what a single-chunk pipeline experiences.
  2. Click the top result to mark it irrelevant, then lower k to 2. Watch hit rate fall to 0 while recall and precision degrade gradually — the binary metric is far more brittle.
  3. Mark every result irrelevant, then mark just the last one relevant. Hit rate stays 0 until k reaches 10, then snaps to 1: no partial credit anywhere along the way.

What to remember

Hit Rate@k is the proportion of queries with at least one relevant result in the top k. It is binary per query, ignores position and quantity, and measures the ceiling on your pipeline — if the evidence never reaches the model, nothing downstream can fix it. Use it as the first number you establish, set k to the number of chunks you actually pass to the generator, and plot it against k to tell a ranking problem from a retrieval problem. Then reach for Recall@k when queries need several documents, and MRR when position matters.

All four of these read the same object: an ordered list of retrieved results, with each one labelled relevant or not by a human or a strong model. They differ only in what they choose to notice about it — which is why quoting one without saying which is close to meaningless, and why the visualisation above shows all four at once.

The labels are the expensive part. A relevance judgement per query-document pair is human work, and an evaluation set of fifty queries with judged results is worth more than any amount of metric sophistication on top of unjudged data. Build the set first.

Choosing k

The k that matters is the number of chunks you actually pass to the model, because that is where the ceiling sits.

Purposek
The ceiling on your current pipelineThe number passed to the model, usually 3–10
Whether reranking would helpA larger k, 20–50
Diagnosing outright failuresThe largest k you can afford to inspect

Reporting hit rate at a k you do not use is a common mistake. Hit rate@100 of 0.99 says almost nothing about a system that passes 5 chunks to the model.

The complementary measurement is worth making at the same time: answer correctness as a function of the number of chunks passed. It usually rises, peaks, and then declines as irrelevant context starts to interfere. That peak is the k to operate at, and it is often smaller than people expect.

Measuring it

def hit_at_k(retrieved_ids, relevant_ids, k):
    return int(bool(set(retrieved_ids[:k]) & set(relevant_ids)))

def hit_rate(eval_set, retriever, k):
    return sum(
        hit_at_k(retriever(q["question"], k=k), q["relevant_ids"], k)
        for q in eval_set
    ) / len(eval_set)

for k in (1, 3, 5, 10, 20):
    print(k, round(hit_rate(eval_set, retrieve, k), 3))

Two caveats on interpretation:

It depends on the labels. If the evaluation set marks one chunk as relevant when three would answer the question, hit rate is understated — the system may have retrieved a perfectly good chunk that was not labelled.

It hides degradation. A system whose relevant results all move from rank 1 to rank 5 has the same hit rate@5 and is materially worse, because models attend more to earlier context. Pair it with MRR.

Which metrics to report together

For a RAG system, a defensible minimum set:

Hit rate or recall@k at the k you pass to the model — the ceiling.

MRR — whether the answer is arriving near the top.

Groundedness — whether the model is using the context.

Correctness — whether the final answer is right.

Refusal rate on unanswerable questions — whether it declines rather than inventing.

Those five cover the two stages and the two failure directions, and they take one evaluation run to produce. Adding nDCG and completeness is worthwhile when comparing closely-matched configurations.

Questions people ask

Is hit rate the same as recall? Identical when each question has one relevant document. Recall counts the share found when there are several.

What is a good hit rate? Corpus-dependent. Above 0.9 at your operating k is a reasonable ambition for a well-tuned pipeline; compare against your own baseline.

Should I report hit rate or MRR? Both. Hit rate is the ceiling and easy to communicate; MRR shows whether the ordering is good.

Does hit rate improve with a reranker? Not at the same k in the retrieval stage — reranking reorders. It improves hit rate at the smaller k finally passed to the model.

What k should I evaluate at? The one you pass to the model, plus a larger one to see whether reranking would help.

Why is my hit rate high but answers poor? Retrieval is fine and generation is not — check groundedness, the prompt, and whether too many chunks are being passed.

Recap in one screen

  • Hit rate@k is the share of queries with at least one relevant result in the top k — binary per query.
  • It equals recall@k when there is one relevant document per question.
  • Crude by design: it ignores position and extra relevant documents, and it communicates well.
  • Read it across several k: a gap between k = 5 and k = 20 is a reranking opportunity.
  • Read the failing queries individually — they cluster around causes the aggregate hides.

Recall check

0 of 3

Say the answer out loud before you reveal it — recalling it is what makes it stick, and rereading it is not.

  1. Without scrolling back — what is the one-line takeaway from this module?

  2. What does this module say about “The definition, and the averaging that hides in it”?

  3. What does this module say about “Why it is the right first metric for RAG”?

Cheat sheet

What is Hit Rate@k?

The bluntest of the ranking metrics, and often the most useful one to start with. Hit Rate@k is 1 if at least one relevant document appears in the top k, and 0 otherwise. It does not care how many relevant documents there were, or where in the top k they landed. Averaged over a query set, it answers a single question: how often does retrieval put something useful in front of the generator?

GEN AI · vizlearn.in/gen_ai/hit_rate_at_k.html

About the author

Ashish Jangra builds and maintains VizLearn. Every module here is written and the visualisation behind it hand-built, so the numbers in a readout come from the same code that draws the picture. Corrections are genuinely welcome and get priority over everything else — if a page states something wrong, or an animation misrepresents what the algorithm does, get in touch.