Completeness in LLM evaluation

Did the answer cover everything the question required? Completeness measures the proportion of the required points that the answer actually states. It is the dimension a confident, correct, perfectly grounded answer can fail outright — and the one users notice last, because a partial answer looks exactly like a full one until you act on it.

Overview

Complete relative to what

Completeness is undefined without a statement of what the answer needed to contain, so the reference is not optional. In practice that means a key-points list per evaluation query: the facts a good answer must include, written by whoever understands the domain.

Writing that list is the real work, and it is worth doing carefully because it is reusable. The same list drives correctness (are the stated facts right?) and completeness (were they all stated?), so one artefact serves two dimensions.

It also forces a decision your users have already made implicitly: is the caveat about digital goods a required point or a nice-to-have? Teams often discover during this exercise that they do not agree on what a good answer is, which is more valuable than any score.

Parameters

Visualisation

—

Readout

What to watch

  • Completeness counts required points covered, not claims made.
  • Adding true, grounded, on-topic claims does not raise it.
  • It is the mirror image of relevance — read them together.

Completeness in LLM evaluation: A Practical Guide

What is completeness, and why can a perfectly correct answer still fail on it?

Why the other dimensions cannot see it

Consider an answer that states only "Refunds are issued within 14 days." when the question was about the full refund policy. It is correct. It is grounded. It is entirely relevant. It scores 1.0 on three dimensions and leaves out the condition that makes it actionable.

This asymmetry is the reason completeness is worth measuring separately. The other three dimensions all penalise saying the wrong thing. Only completeness penalises not saying the right thing, and omission is the harder failure to notice, because there is nothing on the screen to catch your eye.

It is also the failure with the worst consequences in practice. A user who reads a wrong answer may check it. A user who reads a partial answer has no signal that anything is missing and acts on it.

Where incompleteness comes from

Retrieval, most often. If a required point was never in the retrieved context, the model cannot state it without hallucinating. Low completeness alongside low recall@k is a retrieval problem, and no prompt change will fix it. This is the single most common cause.

Chunking. A policy split across two chunks where only one was retrieved gives a confidently half-right answer — which is exactly what parent-document retrieval and generous chunk overlap exist to prevent.

Length limits. A max-tokens cap or a "be concise" instruction trades completeness for brevity, usually without anyone deciding to.

The model stopping early. Given ten chunks, models reliably use the first few more than the rest. A required point in chunk eight is retrieved, in context, and still absent from the answer — which is why MRR is worth tracking next to recall.

Measuring it, and the trap in optimising it

The measurement is mechanical once the key-points list exists: for each required point, ask a judge whether the answer states it, and take the proportion. Per point rather than per answer, so the output tells you which point was dropped — that is the actionable part.

The trap is that completeness is trivially gamed by verbosity. An answer that dumps the entire retrieved context scores 1.0. If completeness is the only dimension you optimise, you will get long, hedged, exhaustive answers that score beautifully and that nobody wants to read.

Which is why it is read against relevance. The pair behaves like precision and recall: completeness is the recall of required information, relevance is its precision, and a system is only good when both are high. Reporting either alone is the same mistake as reporting recall without k.

Did it cover everything the question needed?

Completeness asks whether the answer includes all the information required. It is the counterweight to relevance: an answer can be perfectly on-topic and still leave out something essential.

Question: "What are the eligibility requirements for parental leave?"
Context: "Employees must have 26 weeks' continuous service by the qualifying week, must give 8 weeks' notice, and must provide a MATB1 certificate."
Answer A: "You need 26 weeks' continuous service, 8 weeks' notice and a MATB1 certificate." — complete
Answer B: "You need 26 weeks' continuous service." — correct, grounded, relevant, and incomplete

Answer B passes correctness (what it says is true), groundedness (it is supported) and relevance (it addresses the question). Only completeness catches it — which is why the metric exists as a separate measurement.

Incompleteness is more dangerous than being wrong in one specific way: a partial answer looks authoritative, and the user has no signal that something is missing.

Measuring it

Completeness needs a reference for what "everything" means, and there are two sources.

Against a reference answer. Decompose the reference into required points and check how many appear in the model's answer:

completeness = required points covered / total required points

Against the retrieved context. Check whether the answer used all the relevant information available in the context. This measures the generation stage specifically, without needing hand-written references.

Question: {question}
Reference answer: {reference}
Model answer: {answer}

List the distinct pieces of information in the reference answer. For each,
state whether the model answer covers it.

Reply as JSON: {"points": [{"point": "...", "covered": true|false}]}

Scoring per point rather than per answer is what makes the metric actionable: "3 of 4 required points, missing the notice period" tells you something a single number does not.

The distinction between the two references matters diagnostically. If the answer is incomplete relative to the context, generation is at fault. If the context itself was incomplete, retrieval is — and no prompt change will fix it.

Why answers come out incomplete

The context was incomplete. Retrieval returned two of the three chunks needed. This is the most common cause, and it is a recall problem — particularly for multi-hop questions whose answer spans documents.

The relevant chunk was ranked low. Present at position 9 of 10, where models attend to it less — the "lost in the middle" effect. A reranker addresses this.

The model stopped early. Answering the first part of a multi-part question and treating that as done.

Length constraints. An instruction to be concise, or a low token limit, can truncate required detail. This is a real tension with relevance.

Chunking split the requirements. A list of five conditions cut after the third, so the retrieved chunk looks complete and is not.

That last one is worth watching for, because it is invisible in the answer: the model faithfully reports what it was given, and what it was given was a fragment that reads as whole.

Counting what the answer left out

Completeness is the only dimension that scores what is ABSENT, which is why the others cannot see it. This measures it by decomposing the question, then shows the trap: the metric can be gamed by saying more, and the way out is not a better completeness score.

example_01.pyNumPy
Output

Things to try

  1. Three of the five claims are required points and the reference lists four. Completeness is 0.75 — one required point is simply absent from the answer.
  2. Click the off-topic phone claim on and off. Completeness does not move at all: adding material that was not required cannot improve it.
  3. Compare with relevance on the same claim list. The claims that hurt relevance are exactly the ones completeness ignores, which is why the pair has to be read together.

What to remember

Completeness is the proportion of required points an answer actually states, and it needs a key-points reference to be defined at all. It is the one dimension that penalises omission rather than error, which makes it the failure users notice last and act on first. Low completeness is usually a retrieval or chunking problem rather than a generation one. Measure it per point so you know what was dropped, and always read it against relevance — optimised alone it rewards dumping the entire context into the answer.

These four are usually measured with an LLM as judge: a strong model is shown the question, the retrieved context, the answer and sometimes a reference, and asked to score one dimension at a time. Judging one dimension per call is not a stylistic choice — a single prompt asking for all four produces correlated, mushy scores, because the model settles on a general impression and applies it everywhere.

Two habits make judge scores trustworthy. Ask for a decision and a justification, so a human can audit disagreements. And calibrate against a small human-labelled set before believing any of it — a judge that agrees with your annotators 70% of the time is not measuring what you think it is measuring.

The tension with relevance and concision

Completeness and relevance pull in opposite directions, and pushing either too hard damages the other.

 High completenessLow completeness
High relevanceThe targetTerse, missing detail
Low relevancePadded, buried answerPoor on both

Two ways to have both:

Structure the answer. Direct answer first, then the supporting detail. Nothing is missing, and nothing is buried. Requiring a specific format — a short answer then a bulleted list of conditions — achieves this reliably.

Distinguish required from optional. Instruct the model to include everything the question asked for and nothing it did not. That is a narrower instruction than "be thorough", which produces padding.

The practical target is not maximum completeness. It is sufficient completeness — everything the question needed, and no more.

Multi-hop questions

The case where completeness fails hardest is a question whose answer requires combining two documents.

"Which of our offices are in countries where the standard VAT rate is above 20%?"

That needs the office list and the VAT rates, which are almost certainly in different documents. A single retrieval pass ranked by similarity to the whole question may return one and not the other, because neither document is a good match for the combined query.

Three approaches:

Decompose the question into sub-questions, retrieve for each, and combine the context. The most reliable fix.

Iterative retrieval — retrieve, let the model identify what is still missing, retrieve again. More capable and slower.

Increase k and hope both appear. Cheap, unreliable.

Measuring completeness separately is what makes multi-hop weakness visible. On single-hop factoid sets it looks fine; on multi-hop questions it collapses, and the aggregate hides which.

Practical guidance

Score per point, not per answer. The list of missed points is the actionable output.

Measure against both the reference and the context. The difference localises the failure to retrieval or generation.

Include multi-part and multi-hop questions in the evaluation set, deliberately. A set of single-fact questions will show high completeness and tell you nothing.

Track it alongside relevance and answer length. A completeness improvement that doubled the answer length may not be an improvement.

Watch chunk boundaries for lists and conditions — splitting an enumerated list is a silent completeness failure.

Questions people ask

Is completeness the same as recall? Related but distinct: recall is about retrieved documents, completeness about the generated answer's coverage.

How do I know what "complete" means? From a reference answer, or from the relevant information present in the retrieved context. The second needs no hand-written reference.

Does asking for longer answers improve it? It improves coverage and damages relevance. Structure the answer instead.

Why do multi-part questions fail? The model addresses the first part, and retrieval often only covered that part. Decompose them.

Should a refusal be scored on completeness? Not meaningfully — exclude refusals and track them as their own metric.

What is a good score? Above 0.9 on single-hop questions is achievable; multi-hop is typically much lower and is where the useful work is.

Recap in one screen

  • Completeness asks whether the answer covers everything the question required.
  • A partial answer can pass correctness, groundedness and relevance — only this metric catches it.
  • Score per required point; the list of missing points is what tells you what to fix.
  • Measure against both the reference and the retrieved context to localise the failure.
  • It pulls against relevance — structure the answer rather than lengthening it, and test multi-hop questions deliberately.

Recall check

0 of 3

Say the answer out loud before you reveal it — recalling it is what makes it stick, and rereading it is not.

  1. Without scrolling back — what is the one-line takeaway from this module?

  2. What does this module say about “Complete relative to what”?

  3. What does this module say about “Why the other dimensions cannot see it”?

Cheat sheet

Completeness in LLM evaluation

Did the answer cover everything the question required? Completeness measures the proportion of the required points that the answer actually states. It is the dimension a confident, correct, perfectly grounded answer can fail outright — and the one users notice last, because a partial answer looks exactly like a full one until you act on it.

GEN AI · vizlearn.in/gen_ai/completeness_in_llm_evaluation.html

About the author

Ashish Jangra builds and maintains VizLearn. Every module here is written and the visualisation behind it hand-built, so the numbers in a readout come from the same code that draws the picture. Corrections are genuinely welcome and get priority over everything else — if a page states something wrong, or an animation misrepresents what the algorithm does, get in touch.