Correctness in LLM evaluation

Is the answer right? Correctness compares the answer's claims against a reference answer or against verifiable fact. It is the dimension users care about most and the one that is hardest to automate, because unlike groundedness it cannot be checked against the context — it needs a ground truth that someone has to produce.

Overview

Correct against what, exactly

Correctness is only defined relative to a reference, and choosing that reference is most of the work:

A written gold answer. The usual approach. A human writes the ideal answer per evaluation query and a judge compares. Precise, and expensive enough that evaluation sets stay small — which is why fifty carefully written references beat five hundred sloppy ones.

A key-facts list. Rather than prose, the reference is a set of facts the answer must state. Easier to write, much easier to judge consistently, and it composes naturally with completeness.

Verifiable computation. Where the answer is code, a number or a query, correctness can be executed rather than judged. This is the gold standard when it is available and it almost never is for open-ended questions.

Parameters

Visualisation

—

Readout

What to watch

  • Correctness is judged against a reference, not against the context.
  • A grounded claim can still be wrong if the source document is wrong.
  • Click claims to see correctness and groundedness diverge.

Correctness in LLM evaluation: A Practical Guide

What is correctness, and why is it the most expensive dimension to measure?

Why exact match fails, and what replaces it

The obvious automation — string comparison against the reference — fails immediately on natural language. "14 days", "fourteen days" and "two weeks" are the same answer and share no characters. Exact match systematically punishes fluent phrasing.

The n-gram metrics inherited from translation (BLEU, ROUGE) are only a partial improvement: they reward surface overlap, so a wrong answer phrased like the reference can outscore a right answer phrased differently. They remain in use because they are cheap and deterministic, and they should not be trusted as a primary signal for factual correctness.

What works is claim-level judging: decompose both the answer and the reference into atomic facts, then check each answer claim against the reference. That handles paraphrase, gives partial credit, and tells you which fact was wrong — which a single similarity score never will.

The failure this dimension exists to catch

The important case is the grounded but wrong answer. The model faithfully reports what a retrieved document says, and that document is out of date, contradicted by a newer one, or simply mistaken.

Groundedness scores this perfectly. Relevance scores it perfectly. The answer is fluent, cited, and misinforms the user. No amount of prompt engineering fixes it, because the model did exactly what it was asked; the corpus is the defect.

That makes correctness the dimension that audits your data rather than your model. When correctness is low while groundedness is high, stop looking at the pipeline and go and look at the documents: stale versions, superseded policies, and duplicate chunks disagreeing with each other are the usual culprits.

Judging it without fooling yourself

Show the judge the reference. A model asked "is this answer correct?" with no ground truth is being asked to recall the fact itself, and will confidently score wrong answers as right in exactly the domains where you needed evaluation most.

Ask for partial credit. Binary correct/incorrect throws away the difference between an answer with one wrong detail and one that is entirely wrong. Per-claim scoring gives a proportion.

Calibrate against humans. Have annotators label a sample and measure agreement with your judge. If the judge agrees 70% of the time, a 3-point movement in your correctness score is noise.

Watch for self-preference. A judge tends to favour answers written in its own style, and to prefer longer answers. Both biases are documented and both inflate scores in ways that have nothing to do with being right.

Is the answer right?

Correctness compares the model's answer against a known ground truth. It is the metric stakeholders care about, and the hardest of the RAG metrics to measure well — because natural language has many correct forms.

Question: "How long is parental leave?"
Truth: "39 weeks, of which 6 are at full pay."
Answer A: "You get 39 weeks, with the first 6 paid in full." — correct
Answer B: "Thirty-nine weeks." — correct but incomplete
Answer C: "39 weeks at full pay." — wrong in a detail that matters

Exact string matching marks all three wrong. That is why correctness is usually scored by a model or a human rather than by comparison.

The methods, in ascending order of cost and quality:

MethodSuits
Exact matchShort factual answers, extracted values
Token overlap (F1, ROUGE)Extractive answers
Embedding similarityA loose signal, easily fooled
Model as judgeFree-form answers — the practical default
Human reviewThe reference, and the slowest

Model-as-judge, done carefully

Asking a model to grade another model's answer works well enough to be the standard approach, and it needs care to be trustworthy.

Question: {question}
Reference answer: {truth}
Model answer: {answer}

Does the model answer convey the same information as the reference?
Ignore differences in wording, length and style. Note any factual
contradiction or missing essential detail.

Reply with JSON: {"verdict": "correct" | "partially_correct" | "incorrect",
                  "reason": "<one sentence>"}

Four practices separate a useful judge from a noisy one:

Give it the reference answer. Judging correctness without a ground truth is asking the model to know the answer, which reintroduces the problem you were measuring.

Use a discrete scale. Three or four labels are far more consistent than a 1–10 score, where models cluster around 7 and 8.

Require a reason. It improves consistency and makes disagreements auditable.

Validate against humans on a sample. Grade 50 answers manually and measure agreement with the judge. Below about 80% agreement, the judge's absolute numbers are not usable — though relative comparisons between configurations may still be.

Known biases to guard against: judges prefer longer answers, prefer answers from the same model family, and are sensitive to the order in which candidates are presented. Randomising order and controlling for length help.

Correctness against groundedness

These are different questions and the pair is diagnostic:

GroundedCorrectDiagnosis
YesYesWorking as intended
YesNoRetrieval returned the wrong document, faithfully used
NoYesAnswered from parametric knowledge — unreliable
NoNoHallucination

The third row is the one that looks like success and is not. A system that answers correctly from the model's own memory rather than from the retrieved documents will fail silently as soon as a question moves outside what the model happened to memorise — and no amount of testing on well-known facts will reveal it.

Measuring both is therefore not redundant. Groundedness tells you whether the system is using its sources; correctness tells you whether the result is right.

Why string matching fails, and what the judge costs

Correctness needs a reference answer, and comparing free text to a reference is harder than it sounds. This scores the same answers three ways -- exact match, token overlap, and a judge -- and shows exactly which kinds of answer each one gets wrong.

example_01.pyNumPy
Output

Things to try

  1. Click the digital-goods claim to mark it supported by the context. Groundedness rises to 1.0 and correctness does not move — it is still contradicted by the reference.
  2. That single claim is the grounded-but-wrong case: cited, confident, and misinforming. It is why correctness cannot be inferred from the context.
  3. Compare correctness with completeness. An answer can state only true things and still omit half of what was asked.

What to remember

Correctness compares the answer against a reference or verifiable fact, which makes it the only one of the four dimensions that requires ground truth someone has to write. Exact match and n-gram overlap fail on paraphrase; claim-level judging against a key-facts reference is what works. Its distinctive value is catching the grounded-but-wrong answer, where the model faithfully reports a document that is itself out of date — a defect in your corpus that every other dimension scores as a success.

These four are usually measured with an LLM as judge: a strong model is shown the question, the retrieved context, the answer and sometimes a reference, and asked to score one dimension at a time. Judging one dimension per call is not a stylistic choice — a single prompt asking for all four produces correlated, mushy scores, because the model settles on a general impression and applies it everywhere.

Two habits make judge scores trustworthy. Ask for a decision and a justification, so a human can audit disagreements. And calibrate against a small human-labelled set before believing any of it — a judge that agrees with your annotators 70% of the time is not measuring what you think it is measuring.

Building the evaluation set

Correctness needs ground truth, and where it comes from determines what the number means.

Hand-written by domain experts. Highest quality. 50 questions with careful reference answers is enough to catch regressions; 200 gives stable comparisons.

Derived from existing content. FAQ pairs, resolved support tickets, previously answered queries. Realistic and cheap where they exist.

Generated then reviewed. A model drafts questions and answers from the documents, a human corrects them. The pragmatic route for a new system, and the review step is not optional — unreviewed generated ground truth encodes the generator's errors.

Three things to include deliberately:

Unanswerable questions, where the corpus genuinely does not contain the answer. The correct response is a refusal, and measuring whether the system refuses is as important as measuring whether it answers.

Multi-hop questions requiring two documents. These fail differently and are where most systems are weakest.

Adversarial cases — questions with a false premise, ambiguous questions, questions whose answer changed between document versions.

What correctness does not tell you

Where the failure was. A wrong answer might be a retrieval failure, a prompt failure or a model failure. Correctness alone cannot say, which is why retrieval metrics are measured separately.

Whether the answer was useful. A technically correct answer can be unhelpful — too terse, badly structured, not addressing what was actually asked. Relevance and completeness cover that.

How it behaves on the questions you did not test. A 200-question evaluation set is a sample, and real traffic is broader and stranger.

Whether it is calibrated. A system that is 80% correct and confident every time is more dangerous than one that is 80% correct and flags its uncertainty.

So correctness sits in a suite alongside groundedness, relevance, completeness and refusal rate — not on its own.

Running it as a regression suite

The practical value comes from automation, not from a one-off measurement.

Run it on every change — chunk size, embedding model, prompt wording, model version. Each of those can move correctness, and without a suite the movement is invisible.

Fix the seed and temperature 0 so a difference between runs is a real difference.

Track per-question results, not just the aggregate. A change that fixes ten questions and breaks eight looks like a small improvement and is worth understanding.

Keep a hard subset of questions the system has failed before, as a regression guard.

Version the evaluation set and record which version produced each score, or comparisons across time become meaningless.

Questions people ask

How many questions do I need? 50 to catch large regressions, 200 for stable comparison between similar configurations.

Can I trust a model as judge? For relative comparisons, generally yes. For absolute numbers, validate against human grading on a sample first.

Should I use exact match? Only for short extracted values — dates, amounts, identifiers. It fails on anything free-form.

What about answers that are correct but incomplete? Use a three-level scale, or measure completeness as a separate metric.

How do I handle questions with several valid answers? Provide several reference answers and accept a match against any.

Should the judge be the same model being evaluated? Preferably not — self-preference bias is documented. Use a different model, or a human sample.

Recap in one screen

  • Correctness compares the answer against known ground truth, and exact matching fails on free-form text.
  • Model-as-judge is the practical default: give it the reference, use a discrete scale, require a reason, validate against humans.
  • Correct-but-ungrounded is a warning, not a success — the system answered from memory.
  • Include unanswerable, multi-hop and adversarial questions in the set.
  • Run it as a versioned regression suite at temperature 0, tracking per-question results.

Recall check

0 of 3

Say the answer out loud before you reveal it — recalling it is what makes it stick, and rereading it is not.

  1. Without scrolling back — what is the one-line takeaway from this module?

  2. What does this module say about “Correct against what, exactly”?

  3. What does this module say about “Why exact match fails, and what replaces it”?

Cheat sheet

Correctness in LLM evaluation

Is the answer right? Correctness compares the answer's claims against a reference answer or against verifiable fact. It is the dimension users care about most and the one that is hardest to automate, because unlike groundedness it cannot be checked against the context — it needs a ground truth that someone has to produce.

GEN AI · vizlearn.in/gen_ai/correctness_in_llm_evaluation.html

About the author

Ashish Jangra builds and maintains VizLearn. Every module here is written and the visualisation behind it hand-built, so the numbers in a readout come from the same code that draws the picture. Corrections are genuinely welcome and get priority over everything else — if a page states something wrong, or an animation misrepresents what the algorithm does, get in touch.