Relevance in LLM evaluation
Does the answer address the question that was asked? Relevance measures how much of the answer is on point — and it is the dimension that catches the failure the other three are blind to: a response that is entirely true, fully grounded, and quietly about something else.
Overview
Two different things are called relevance
The word is used for two distinct measurements and conflating them makes evaluation results incomparable.
Context relevance scores the retrieved chunks against the query. It is a retrieval metric, closely related to precision@k, and it tells you whether your retriever is sending noise to the generator.
Answer relevance scores the generated answer against the query. It is a generation metric, and it is what this page is about.
They fail independently. Perfect context and an irrelevant answer means the generator wandered; irrelevant context and a relevant answer usually means the model answered from memory, which groundedness will catch. Always say which one you mean.
Parameters
Visualisation
—Readout
What to watch
- The 24/7 phone claim is true and grounded and still off topic.
- Padding is scored here and nowhere else.
- Click claims to see relevance move independently of the rest.
Relevance in LLM evaluation: A Practical Guide
What is answer relevance, and how does it differ from context relevance?
The failures it exists to catch
Padding. The answer contains the requested information plus three paragraphs of adjacent context nobody asked for. Every claim is true and grounded, and the user has to hunt for the answer. Models trained to be helpful pad heavily, and no other dimension penalises it.
Answering a nearby question. Asked how long refunds take, the answer explains how to request one. Fluent, grounded, correct, and not responsive.
Hedging. Several paragraphs of caveats and conditions with no actual answer inside them. Technically nothing is wrong; nothing is useful either.
Restating the question. The degenerate answer that echoes the query back. It scores oddly well on naive similarity-based relevance measures, which is a good reason not to use them.
How to measure it without measuring similarity
The tempting approach — embed the question and the answer and take cosine similarity — is bad. It rewards vocabulary overlap, so restating the question scores highly and a correct answer that shares no words with the question scores poorly. "When are refunds issued?" answered with "Within a fortnight of delivery" is perfect and lexically distant.
The approach that works is question generation: ask a model to write the questions this answer would be a good response to, then compare those against the real question. If the generated questions match, the answer is on point; if they are broader or different, it is padded or off target. It measures what relevance actually means rather than a proxy for it.
The alternative is claim-level judging, as in the visualisation above: split the answer into claims and mark each as responsive or not. Relevance is then the supported proportion, and you can see exactly which sentence was padding.
The tension with completeness, and how to resolve it
Relevance and completeness pull in opposite directions, and optimising either alone produces a bad system.
Maximise relevance alone and answers get terse to the point of being unhelpful — every qualification stripped out because qualifications are not strictly what was asked. Maximise completeness alone and answers sprawl, because adding material can only help.
The resolution is to score both and watch them together, the same way precision and recall are read as a pair. An answer that is 0.95 relevant and 0.6 complete is leaving things out; one that is 0.6 relevant and 1.0 complete is burying the answer in padding. Neither single number would tell you which problem you have.
Note also that relevance is the dimension most affected by your prompt rather than your retrieval. "Answer in at most three sentences, addressing only what was asked" moves it substantially, and costs nothing.
Did it answer the question that was asked?
Relevance asks whether the response addresses the user's actual question. It is separate from correctness, and a system can fail on it while being entirely factual.
Question: "How do I cancel my subscription?"
Answer A: "Go to Settings, then Billing, then Cancel." — relevant
Answer B: "Our subscription plans include Basic at £9 and Pro at £29." — true, and not an answer
Answer C: "Cancellation is available. You may also downgrade, pause, transfer or gift your subscription, and our refund policy states..." — buried|
B is the classic retrieval-driven failure: the model was given documents about subscriptions and produced something about subscriptions. C is the padding failure, which is common in tuned models and genuinely reduces usefulness.
Three distinguishable problems hide under low relevance:
| Problem | Symptom |
|---|---|
| Off-topic | Answers a different question |
| Partial | Addresses one part of a multi-part question |
| Padded | Correct answer buried in unrequested material |
Measuring it
Relevance has no ground-truth answer to compare against — it is judged against the question, not against a reference. That makes it a natural fit for a model judge, and it means the judging prompt has to be precise about what is being asked.
Question: {question}
Answer: {answer}
Does the answer directly address what was asked? Ignore whether it is
factually correct - judge only whether it responds to the question.
Reply as JSON: {"verdict": "relevant" | "partially_relevant" | "irrelevant",
"reason": "<one sentence>"}The instruction to ignore correctness matters. Without it, judges conflate the two and the metric stops being independently useful.
An alternative measurement worth knowing: answer-to-question similarity via generated questions. Ask a model what question the answer appears to be answering, embed both that and the original question, and compare. A large gap indicates the answer drifted. It is cheaper than a judge and cruder.
Why answers drift off-topic
Diagnosing low relevance means knowing the causes, and most of them are not the model's fault.
Retrieval returned topically-related but unhelpful chunks. The model uses what it is given, so context about subscription pricing produces an answer about pricing. This is the most common cause, and it is a retrieval problem.
The question was ambiguous. "How do I change it?" needs the conversation history resolved before retrieval; without contextualisation, both retrieval and generation guess.
Multi-part questions get partially answered, because the model addresses the first part and the retrieved context only covers that part.
The prompt encourages padding. Instructions to be "helpful and thorough" reliably produce answers with unrequested extras.
Instruction tuning bias. Models tuned on preference data tend towards longer answers, because human raters mildly prefer them. That is a systematic pull towards padding.
Measuring "answers the question" without measuring similarity
The obvious way to score relevance is to compare the answer with the question, and it is the wrong way -- a restatement of the question is maximally similar to it and answers nothing. This shows that failure directly, then the two things the dimension is actually for.
Things to try
- The 24/7 phone claim is marked correct but off topic. Relevance drops while correctness and groundedness stay put — this is the padding case, visible only here.
- Click that claim to flip its grounding. Groundedness moves and relevance does not: whether the context supports a claim has nothing to do with whether it answers the question.
- Compare relevance and completeness across the claim list. The off-topic claim hurts one and does nothing for the other, which is why they have to be read as a pair.
What to remember
Answer relevance measures how much of the answer addresses the question asked. Distinguish it from context relevance, which scores retrieved chunks and is a retrieval metric. It is the only dimension that penalises padding, hedging and confidently answering a nearby question — all of which score perfectly on correctness and groundedness. Do not measure it with question-answer embedding similarity, which rewards restating the question; use generated questions or per-claim judging. Read it alongside completeness, because optimising either alone makes answers worse.
These four are usually measured with an LLM as judge: a strong model is shown the question, the retrieved context, the answer and sometimes a reference, and asked to score one dimension at a time. Judging one dimension per call is not a stylistic choice — a single prompt asking for all four produces correlated, mushy scores, because the model settles on a general impression and applies it everywhere.
Two habits make judge scores trustworthy. Ask for a decision and a justification, so a human can audit disagreements. And calibrate against a small human-labelled set before believing any of it — a judge that agrees with your annotators 70% of the time is not measuring what you think it is measuring.
What improves it
Fix retrieval first. If the retrieved chunks are about the right topic and the wrong question, the answer will be too. Measure recall@k on the specific failing queries before touching the prompt.
Contextualise the question against conversation history before retrieving. A follow-up question is meaningless as a standalone search, and this single step fixes a large share of relevance failures in chat interfaces.
Decompose multi-part questions, retrieve for each part, and instruct the model to address each explicitly.
Constrain the shape of the answer. "Answer in at most three sentences. Do not include information that was not asked for." Blunt, and effective against padding.
Ask for the direct answer first, then any caveats. This helps even when the extra material stays, because the useful part is no longer buried.
Reduce the number of chunks. Ten chunks invite the model to synthesise across all of them; three keep it focused.
Relevance within the metric suite
| Metric | Question | Independent of |
|---|---|---|
| Recall@k | Was the answer retrieved? | Generation |
| Groundedness | Is each claim supported by the context? | Truth |
| Correctness | Is the answer right? | Style |
| Relevance | Does it address the question? | Truth and completeness |
| Completeness | Does it cover everything asked? | Concision |
Relevance and completeness pull against each other, which is worth naming explicitly. A terse answer scores well on relevance and may miss required detail; an exhaustive one covers everything and buries the answer.
Reporting both, and reading them as a pair, is how you find the balance. High relevance with low completeness means answers are too narrow; the reverse means they are padded.
Practical evaluation setup
Three things make a relevance measurement trustworthy:
Use real user questions. Generated questions tend to be well-formed and single-part, which is exactly the case relevance failures do not occur in. Mine logs where possible.
Include multi-part and ambiguous questions deliberately. These are where relevance breaks, and a set of clean factoid questions will show 0.95 and tell you nothing.
Validate the judge on a sample. Grade 50 answers yourself and measure agreement. Judges are more consistent on relevance than on correctness, because the question is narrower — and it is still worth checking.
Track it per question type as well as in aggregate. Relevance is frequently high on single-part factoid questions and much lower on comparative or procedural ones, and the aggregate hides that.
Questions people ask
Is relevance the same as correctness? No. An answer can be entirely true and not address the question.
How is it different from precision@k? Precision measures retrieved documents; relevance measures the generated answer.
Can I measure it without ground truth? Yes — it is judged against the question, which is what makes it cheaper to evaluate than correctness.
Why are my answers so long? Instruction-tuned models are biased towards length, because preference raters mildly prefer it. Constrain it explicitly in the prompt.
Should a refusal count as relevant? A well-targeted refusal — "the documents do not cover this" — is relevant. Track refusal rate separately so over-refusal is visible.
What if the question itself is unclear? Asking a clarifying question is arguably the most relevant response. Decide whether your evaluation set rewards that.
Recap in one screen
- Relevance asks whether the answer addresses the question, independently of whether it is true.
- Three failure shapes: off-topic, partial, and correct-but-buried.
- The most common cause is retrieval returning topically-related but unhelpful chunks.
- Contextualise follow-up questions, decompose multi-part ones, and constrain answer length.
- Read it alongside completeness — the two pull in opposite directions.