Groundedness in LLM evaluation

Every claim in the answer should trace back to something in the retrieved context. Groundedness measures how much of the answer is actually supported by the evidence the model was given. It is the direct, automatable measure of hallucination — and it is deliberately blind to whether those claims are true, which is what separates it from correctness.

Overview

What it measures, claim by claim

The answer is decomposed into atomic claims — individual assertions that could each be checked independently — and each is tested against the retrieved context. Groundedness is the proportion that are supported.

The decomposition matters. Scoring a whole paragraph as "grounded or not" throws away the most useful information a groundedness evaluation produces: which sentence was invented. A per-claim score points you at the exact span to investigate, and lets you distinguish an answer that is 90% supported with one fabricated detail from one that is wholly invented.

"Supported" should mean entailed by the context, not merely consistent with it. A claim the context neither states nor contradicts is ungrounded, even if it is plausible and even if it is true. That strictness is the whole value of the metric.

Parameters

Visualisation

—

Readout

What to watch

  • Click a claim to flip whether the context supports it.
  • A claim can be true and ungrounded — the model knew it, the context did not say it.
  • Groundedness needs no reference answer, only the context.

Groundedness in LLM evaluation: A Practical Guide

What is groundedness, and how is it different from correctness?

Why it is independent of correctness

This is the distinction people collapse, and the four combinations are all real:

Grounded and correct. The context said it and it is true. What you want.

Grounded and wrong. The context said it and the context is out of date or simply incorrect. The model behaved perfectly; your corpus is the problem. Groundedness scores 1.0 and the user is misinformed.

Ungrounded and correct. The model knew the answer from pretraining and the context never mentioned it. Harmless-looking, and it is the failure mode that erodes trust in a RAG system: you have no idea which answers came from your documents and which came from the model's memory, so the citations mean nothing.

Ungrounded and wrong. A hallucination, in the ordinary sense.

Groundedness catches rows three and four without needing to know the truth of anything. That is why it is cheap: it requires the context and the answer, and no reference and no domain expert.

How it is measured in practice

LLM-as-judge, per claim. Split the answer into claims, then for each ask a strong model whether the context entails it, with the answer justified. This is the standard approach and it works well, because entailment against a supplied passage is a much easier task than open-ended judgement.

Natural-language inference models. A dedicated NLI classifier scores entailment between context and claim. Cheaper and faster than a large judge, weaker on long or technical passages.

Citation checking. Require the model to cite a chunk id per sentence, then verify the cited chunk actually supports it. This has the useful property of being auditable by a human in seconds, and it changes the generator's behaviour for the better even before you measure anything.

What a low score is actually telling you

A groundedness problem is often a retrieval problem wearing a generation costume. If the context did not contain the answer, a model asked to answer anyway will fill the gap from memory — so low groundedness correlates with low recall@k, and the fix is upstream.

The other common cause is a prompt that does not permit abstention. A model told to "answer the question using the context" will answer; one told "if the context does not contain the answer, say so" will often decline correctly. That single sentence moves groundedness more than most model changes, and it is why corrective RAG treats declining as a first-class outcome.

Watch for the degenerate optimum, too. An answer that quotes the context verbatim and says nothing else scores 1.0 on groundedness and may be useless. Groundedness has to be read next to completeness and relevance, or you will optimise your way into a system that recites.

Is every claim supported by the sources?

Groundedness — also called faithfulness — asks a narrow question: does each statement in the answer follow from the retrieved context?

It deliberately does not ask whether the answer is true. Those are separate measurements, and separating them is what makes the pair diagnostic.

Context: "Parental leave is 39 weeks. The first 6 weeks are paid at 90% of salary."
Answer A: "You get 39 weeks, with the first 6 at 90% pay." — grounded
Answer B: "You get 39 weeks, with the first 6 at full pay." — not grounded (the context says 90%)
Answer C: "You get 39 weeks. You must give 8 weeks' notice." — the second sentence is unsupported, even if true

Answer C is the case that matters most in practice. The claim may well be correct, and it did not come from the provided documents — so the citation is misleading and the system will produce similar unsupported claims where they happen to be wrong.

Measuring it, claim by claim

Groundedness is scored per claim rather than per answer, because an answer is usually part supported and part not.

The standard procedure:

  1. Decompose the answer into atomic claims.
  2. Check each claim against the retrieved context.
  3. Score as the share of claims that are supported.
Context:
{retrieved_chunks}

Claim: {claim}

Is this claim fully supported by the context above? Answer "supported",
"contradicted" or "not_found", with the sentence from the context that
supports it if applicable. Reply as JSON.

groundedness = supported claims / total claims

The decomposition step is where the difficulty lies. "Parental leave is 39 weeks, of which the first 6 are at 90% pay, and it must be requested 8 weeks in advance" is three claims, and scoring it as one loses the information that two are supported and one is not.

Frameworks such as RAGAS implement this pipeline with a model as judge. The three-way verdict — supported, contradicted, not found — is more useful than a binary one, because contradiction and absence have different causes: contradiction usually means the model mis-read the context, absence means it drew on its own knowledge.

The four-way diagnosis

Groundedness is most useful read against correctness:

GroundedCorrectWhat happenedWhat to do
YesYesWorkingNothing
YesNoWrong document retrieved, faithfully usedFix retrieval
NoYesAnswered from the model's own knowledgeTighten the prompt — this will fail silently
NoNoHallucinationFix retrieval and the prompt

The third row deserves emphasis because it presents as success. A system answering correctly from parametric memory rather than from its sources passes any correctness test built on well-known facts, and fails the moment a question concerns something specific to your organisation — which is the entire reason the system exists.

Measuring groundedness is how you detect that before it reaches users.

Scoring an answer against its sources, claim by claim

Groundedness asks whether every statement in an answer is supported by the retrieved context, and it is the one RAG metric that can be checked without knowing the right answer. This runs the claim-level procedure and shows why the sentence-level shortcut disagrees with it.

example_01.pyNumPy
Output

Things to try

  1. Click the digital-goods claim, which the context does not support. Groundedness falls while correctness is unaffected — the two dimensions move independently.
  2. Flip every claim to supported. Groundedness reaches 1.0 while relevance and completeness stay where they were: a perfectly grounded answer can still be padded and incomplete.
  3. Note that the 24/7 phone claim is marked correct but off topic. True, plausible, and not what was asked — groundedness alone would not catch it.

What to remember

Groundedness is the proportion of an answer's claims that the retrieved context actually supports. It is the direct measure of hallucination, it needs no reference answer, and it is independent of truth — a claim can be grounded and wrong, or true and ungrounded. Measure it per claim rather than per answer so it tells you which span to look at. A low score usually means a retrieval failure or a prompt that does not allow the model to decline. Never optimise it alone, because verbatim quotation scores perfectly.

These four are usually measured with an LLM as judge: a strong model is shown the question, the retrieved context, the answer and sometimes a reference, and asked to score one dimension at a time. Judging one dimension per call is not a stylistic choice — a single prompt asking for all four produces correlated, mushy scores, because the model settles on a general impression and applies it everywhere.

Two habits make judge scores trustworthy. Ask for a decision and a justification, so a human can audit disagreements. And calibrate against a small human-labelled set before believing any of it — a judge that agrees with your annotators 70% of the time is not measuring what you think it is measuring.

What improves it

Better retrieval, first. Most ungrounded answers are the model doing its best with context that did not contain the answer. If recall@k is 0.6, groundedness cannot be high — and prompt work will not fix it.

An explicit refusal option. A model with no permission to say "not in the documents" will always produce something. One sentence in the prompt removes a large class of unsupported claims.

Required citations. Asking for a source id per claim makes unsupported statements structurally awkward and independently verifiable.

Fewer, better chunks. Contradictory or irrelevant context actively causes unsupported claims. Three relevant chunks beat ten containing three.

Temperature 0. Sampling variation in factual output is error.

Structured output. Requiring JSON with a source field per claim makes the constraint explicit.

The prompt that does most of this work:

Answer using only the context below. Cite the source number after each
claim. If the context does not contain the answer, reply exactly:
"Not found in the provided documents."

The judge's limitations

Groundedness scored by a model has known weaknesses, and knowing them determines how much weight to put on the number.

Implicit inference. If the context says "39 weeks" and the answer says "just under 10 months", is that supported? A strict judge says no; a reasonable reader says yes. Judges are inconsistent on paraphrase and arithmetic.

Multi-hop support. A claim following from two chunks combined may be marked unsupported by a judge examining chunks individually.

Decomposition quality. Badly split claims produce meaningless verdicts.

Judge bias. Judges tend to be lenient towards fluent text and towards answers from the same model family.

The mitigation is the same as for any model-graded metric: validate against human judgement on a sample of 50, and if agreement is below about 80%, treat the absolute number as unusable while relative comparisons between configurations may still hold.

Where it sits in the metric suite

MetricQuestion
Recall@kWas the answer retrieved?
GroundednessIs every claim supported by what was retrieved?
CorrectnessIs the answer right?
RelevanceDoes it address the question?
CompletenessDoes it cover everything asked?
Refusal rateDoes it decline when it should?

Groundedness is the one that specifically measures whether the retrieval-augmented part of the system is doing anything. A high correctness score with low groundedness means you have built an expensive pipeline whose documents the model is ignoring.

Questions people ask

Is groundedness the same as correctness? No. Groundedness asks whether claims follow from the context; correctness asks whether they are true. Both are needed.

Can an answer be grounded but wrong? Yes — if the retrieved document is wrong or out of date. That is a retrieval or corpus problem, not a generation one.

How do I decompose an answer into claims? With a model, prompted to split into atomic factual statements. It is imperfect and works well enough in practice.

Should a refusal count as grounded? Yes — there are no unsupported claims. Track refusal rate separately so over-refusal is visible.

What score should I aim for? Above 0.9 for systems where citations matter. Below 0.7 indicates the model is regularly drawing on its own knowledge.

Does it detect subtle numeric errors? Inconsistently — "39 weeks" against "40 weeks" is often caught, unit and percentage changes less so. Worth spot-checking numbers separately.

Recap in one screen

  • Groundedness asks whether every claim follows from the retrieved context, not whether it is true.
  • Score it per claim: decompose the answer, check each against the context, take the supported share.
  • Correct-but-ungrounded means the model answered from memory — it looks like success and fails silently.
  • Better retrieval, an explicit refusal option and required citations are what move it.
  • Validate the model judge against human grading before trusting absolute numbers.

Recall check

0 of 3

Say the answer out loud before you reveal it — recalling it is what makes it stick, and rereading it is not.

  1. Without scrolling back — what is the one-line takeaway from this module?

  2. What does this module say about “What it measures, claim by claim”?

  3. What does this module say about “Why it is independent of correctness”?

Cheat sheet

Groundedness in LLM evaluation

Every claim in the answer should trace back to something in the retrieved context. Groundedness measures how much of the answer is actually supported by the evidence the model was given. It is the direct, automatable measure of hallucination — and it is deliberately blind to whether those claims are true, which is what separates it from correctness.

GEN AI · vizlearn.in/gen_ai/groundedness_in_llm_evaluation.html

About the author

Ashish Jangra builds and maintains VizLearn. Every module here is written and the visualisation behind it hand-built, so the numbers in a readout come from the same code that draws the picture. Corrections are genuinely welcome and get priority over everything else — if a page states something wrong, or an animation misrepresents what the algorithm does, get in touch.