What is MRR (Mean Reciprocal Rank)?
For one query, find the rank of the first relevant result and take its reciprocal: rank 1 scores 1.0, rank 2 scores 0.5, rank 5 scores 0.2, nothing relevant scores 0. Average that across your queries and you have MRR. Everything after the first hit is ignored entirely — which is the point, and the limitation.
Overview
The definition, and the shape of the curve
Reciprocal rank for a query is 1/(rank of the first relevant result). MRR is the mean of that over an evaluation set. The name is worth reading literally: it is a mean of reciprocal ranks, and each word matters.
The reciprocal makes the curve steep at the top and flat at the bottom. Moving a result from rank 2 to rank 1 gains 0.5. Moving one from rank 10 to rank 9 gains about 0.011 — forty-five times less. That shape encodes an assumption: users and models care enormously about the top of the list and barely at all about the bottom.
Because it is a mean of per-query scores, a single query can move it noticeably on a small evaluation set. Report the number of queries alongside it, and be suspicious of MRR differences on fewer than a hundred.
Parameters
Visualisation
—Readout
What to watch
- Only the first relevant result contributes; the rest are ignored.
- The reciprocal drops steeply: rank 1 to rank 2 halves the score.
- Click the top result on and off to see the whole metric swing.
What is MRR (Mean Reciprocal Rank)?: A Practical Guide
What does MRR measure, and when is the position of the first correct result the thing that matters?
When position is the whole question
MRR is the right metric when the consumer stops at the first good result. Question answering with a single correct answer, "I'm feeling lucky" search, entity lookup, a code assistant jumping to a definition — in all of these the second correct result is worth nothing.
It also matters inside RAG more than it first appears. Attention over a long context is uneven, and evidence placed first is more likely to be used than evidence placed eighth. Two retrievers with identical recall@5 can produce measurably different answers if one puts the key chunk at the top and the other buries it. Recall cannot see that difference; MRR can.
The practical use: track recall@k to know whether the evidence arrives, and MRR to know whether it arrives somewhere the model will actually look. A reranker that leaves recall unchanged while lifting MRR is doing real work.
What it ignores, and when that is wrong
Every relevant result after the first. A query with one relevant document at rank 1 and a query with ten relevant documents at ranks 1 to 10 both score 1.0. If your questions need multiple sources — comparisons, summaries, anything aggregative — MRR is close to blind to what you care about, and Recall@k is the metric to use.
Degrees of relevance. A marginally useful document at rank 1 outscores a perfect one at rank 2. MAP averages precision over every relevant position and nDCG additionally uses graded judgements; both are strictly more informative and both cost more to label.
The zero. Queries with no relevant result contribute 0, which is correct and worth watching: a system with excellent MRR on the queries it answers and a large silent tail of zeros has a coverage problem that the average partially conceals. Report the proportion of zero-score queries next to it.
Reading MRR against the alternatives
The family is easiest to keep straight by what each one uses:
Hit Rate@k — is there anything relevant in the top k? Binary, ignores position entirely.
MRR — where is the first relevant result? Uses position, ignores everything after it.
MAP — where is every relevant result? Uses all positions, still binary relevance.
nDCG — where is every relevant result, and how relevant is each? Uses positions and grades, and is the most informative and most expensive.
They are a ladder of increasing information and increasing labelling cost. Climb it only as far as your judgements can honestly support: nDCG computed from binary labels guessed by a weak model is not better than a hit rate from careful human ones.
How high was the first right answer?
MRR measures where the first relevant result appeared, averaged over queries.
MRR = (1/|Q|) Σ 1 / rank of the first relevant result
The reciprocal is what gives the metric its character:
| First relevant at rank | Reciprocal rank |
|---|---|
| 1 | 1.00 |
| 2 | 0.50 |
| 3 | 0.33 |
| 5 | 0.20 |
| 10 | 0.10 |
| Not found in top k | 0 |
Worked example over four queries, with the first relevant result at ranks 1, 3, 2 and not-found:
MRR = (1.00 + 0.33 + 0.50 + 0) / 4 = 0.46
Note how harsh the drop from rank 1 to rank 2 is — half the credit for one position. That is deliberate, and it matches two realities: users scan from the top, and language models attend most strongly to the earliest context.
What it is and is not good for
MRR suits questions with essentially one right answer. Factoid question answering, entity lookup, "which document covers X" — cases where finding the answer once is what matters and there is nothing to gain from finding it twice.
MRR ignores everything after the first hit. A system that puts one relevant document at rank 1 and nine irrelevant ones below scores the same as one that puts ten relevant documents in the top ten. If coverage matters, MRR is the wrong metric.
That makes it complementary rather than competing:
| Metric | Answers |
|---|---|
| Recall@k | Was the answer retrieved at all? |
| MRR | How high was it? |
| Precision@k | How much of what we returned was useful? |
| nDCG@k | How good is the whole ranking, with graded relevance? |
The pair worth reporting together for RAG is recall@k and MRR. Recall says whether the ceiling is high enough; MRR says whether the ordering is putting the answer where the model will attend to it.
Reading the two together
| Recall@10 | MRR | Diagnosis |
|---|---|---|
| 0.95 | 0.85 | Healthy — found and ranked well |
| 0.95 | 0.40 | Found, but buried — add a reranker |
| 0.60 | 0.55 | Ranking is fine; retrieval is missing documents |
| 0.60 | 0.20 | Both stages need work |
The second row is the common and fixable one. High recall with low MRR means the candidate set is good and the ordering is not, which is exactly what a cross-encoder reranker addresses — and it typically moves MRR substantially while leaving recall unchanged.
The third row points the other way: no reranker can help, and the work is in chunking, the embedding model, or adding keyword search.
Scoring by where the first good answer landed
The sections above describe what MRR measures and what it ignores. Here it is as arithmetic on a concrete result set, along with the two numeric traps the description cannot show you: what the zeros do to the average, and why the mean of reciprocal ranks behaves nothing like the reciprocal of the mean rank.
Things to try
- Click result #1 on and off. MRR swings between 1.0 and 0.33 on a single change, while recall moves a fraction — MRR is by far the most position-sensitive of the four.
- Mark only the last result relevant. RR is 0.1, and no amount of additional relevant documents further down would improve it.
- Mark results 1 and 2 both relevant, then mark only result 1. MRR is identical either way — everything after the first hit is invisible to it.
What to remember
MRR averages 1/rank of the first relevant result. It is the metric for systems where the consumer stops at the first good answer, and it is far more sensitive to the top of the ranking than recall or precision. It ignores every relevant result after the first, so it is the wrong choice for questions needing several sources. Inside RAG it is a useful companion to recall: recall says the evidence arrived, MRR says whether it arrived where the model will actually attend to it.
All four of these read the same object: an ordered list of retrieved results, with each one labelled relevant or not by a human or a strong model. They differ only in what they choose to notice about it — which is why quoting one without saying which is close to meaningless, and why the visualisation above shows all four at once.
The labels are the expensive part. A relevance judgement per query-document pair is human work, and an evaluation set of fifty queries with judged results is worth more than any amount of metric sophistication on top of unjudged data. Build the set first.
Measuring it
def reciprocal_rank(retrieved_ids, relevant_ids):
relevant = set(relevant_ids)
for i, doc_id in enumerate(retrieved_ids, start=1):
if doc_id in relevant:
return 1.0 / i
return 0.0
def mrr(eval_set, retriever, k=10):
return sum(
reciprocal_rank(retriever(q["question"], k=k), q["relevant_ids"])
for q in eval_set
) / len(eval_set)Two details that affect the number:
The cutoff matters. MRR@10 treats a relevant document at rank 15 as not found, scoring 0. Report the cutoff you used, and keep it consistent when comparing.
Zeros dominate. A query where nothing relevant was retrieved contributes 0, and with 20% such queries the maximum achievable MRR is 0.8. Low MRR can therefore be a recall problem in disguise — which is why the two are read together.
For per-query debugging, the reciprocal ranks are more useful than the mean. Sort queries by reciprocal rank and look at the zeros first: those are the outright failures, and they usually cluster around a recognisable cause.
MRR and RRF are not the same thing
The similarity in name and formula causes confusion, so it is worth separating them.
MRR is an evaluation metric — 1/rank, averaged over queries, measuring how well a system ranks.
Reciprocal rank fusion is a combination method — Σ 1/(k + rank) summed across several ranked lists, producing a merged ranking.
They share the reciprocal-of-rank idea and do entirely different jobs: one scores a system, the other builds one. The +k in RRF (usually 60) is a smoothing constant that flattens the difference between top ranks; MRR has no such constant, which is why it is so harsh on rank 2.
Related rank-sensitive metrics
MAP (mean average precision) averages the precision at each rank where a relevant document appears, then averages over queries. Unlike MRR it accounts for all relevant documents, which makes it the better choice when several exist.
nDCG@k discounts gains logarithmically by position and supports graded relevance — highly relevant, somewhat relevant, irrelevant. The most complete measure, and it needs graded judgements, which are more expensive to collect.
Hit rate@k is the crude version: did any relevant document appear in the top k, yes or no. Useful as a headline, and blind to position.
For most RAG work, recall@k plus MRR is enough. Move to nDCG when comparing closely-matched configurations and you have graded labels.
Questions people ask
What is a good MRR? Corpus-dependent. Track it against your own baseline over time rather than an absolute target.
Should I use MRR or MAP? MRR when there is one right answer; MAP when several relevant documents exist and all of them matter.
Why is MRR so harsh on rank 2? Because 1/2 is half of 1/1. It reflects that being first matters disproportionately, for users and for model attention.
Does MRR need graded relevance? No — binary judgements are sufficient, which is part of its appeal.
How does a reranker affect it? Substantially, and that is precisely what a reranker is for: reordering a fixed candidate set.
Is MRR the same as reciprocal rank fusion? No — one evaluates a ranking, the other combines several rankings.
Recap in one screen
- MRR averages
1/rankof the first relevant result, so being first is worth twice being second. - It suits single-answer questions and ignores everything after the first hit.
- Report it with recall@k: high recall and low MRR means add a reranker; low recall means fix retrieval first.
- Queries with nothing relevant contribute 0, so low MRR can be a recall problem in disguise.
- MRR evaluates a ranking; reciprocal rank fusion combines rankings — similar formula, different job.