Corrective RAG (CRAG)
Naive RAG retrieves, then generates, whatever came back. Retrieval always returns something, so a query with no good match still produces a confident answer built on irrelevant text. Corrective RAG adds a grader between the two steps: judge the evidence first, and when it is poor, rewrite the query, fall back to another source, or decline to answer.
Overview
The failure it fixes
A vector search returns the k nearest chunks whether or not any of them is relevant. There is no null result: ask about something absent from the corpus and you still get five chunks, at low similarity, and the generator dutifully writes an answer from them.
That is the most damaging RAG failure in production, because the answer is fluent, sourced-looking and wrong. The user has no signal that the retrieval missed.
Parameters
Visualisation
—Readout
What to watch
- Retrieval never fails loudly — it returns the nearest thing it has.
- The grader runs before generation, which is the whole design.
- "I don't know" is one of the paths, not a failure of the system.
Corrective RAG (CRAG): A Practical Guide
What is Corrective RAG, and how does it stop a model answering from irrelevant documents?
The three verdicts
A grader — a small model, a cross-encoder, or a similarity threshold — labels the retrieved set:
Correct. At least one document is relevant. Proceed, often after stripping the irrelevant ones so they do not distract the generator.
Incorrect. Nothing is relevant. Discard everything and take a corrective action rather than generating.
Ambiguous. Something is partly relevant. Combine what you have with an external source.
What correction actually means
Query rewriting is the cheapest: the user's phrasing may simply not match the corpus vocabulary. See query rewriting and HyDE.
Fallback to another source — web search, a second index, a keyword search when the dense one missed.
Abstain. The one people leave out, and often the correct behaviour. A system that answers "that is not in the documents I have" is more useful than one that invents a citation.
The cost is latency and tokens: every query pays for grading, and corrected queries pay for a second retrieval. Grade cheaply, and consider skipping it when the top similarity is already unambiguous.
What the grader can actually be
A similarity threshold. Free, and the crudest: reject when the top score is below a cutoff. It works badly on its own because embedding similarity is not calibrated — 0.7 means different things for different queries, so a fixed threshold rejects good evidence for some queries and accepts noise for others.
A cross-encoder. The reranker you may already run scores query-document pairs jointly and is far better calibrated than a bi-encoder similarity. If a reranker is in the pipeline, its score is the natural grader and costs nothing extra.
A small LLM. "Does this document help answer this question? yes/no." Most accurate, most expensive, and the latency lands on every query rather than only the corrected ones.
The usual production shape is a cheap filter that rejects the obviously bad, with a model consulted only in the ambiguous band.
The costs, and when to skip the grader entirely
Grading is not free. Every query pays for it, and corrected queries pay for a second retrieval and sometimes a second generation — so a pipeline that corrects aggressively can double its p95 latency.
Two mitigations are worth knowing. Skip when unambiguous: if the top result scores far above anything else, grading tells you nothing you did not already know, so gate it on a margin rather than running it always. Cap the retries: one rewrite, then fall back. Without a depth limit, a query the rewriter cannot fix loops until something times out, and each iteration costs a retrieval and a model call.
Measure the benefit rather than assuming it. The number that matters is how often the corrective path fires and how often it then produces a better answer — on a corpus with good coverage, that can be rare enough that the latency is not worth paying.
Where it sits among the adaptive patterns
Corrective RAG is one of a family that all add a decision point to the straight-line retrieve-then-generate pipeline, and they are easy to confuse.
Self-RAG trains the model to emit retrieval and critique tokens itself, so the decision to retrieve and the judgement of what came back are part of generation rather than a separate step. More elegant, and it needs a fine-tuned model.
Adaptive RAG decides whether to retrieve at all based on the query. "What is 2+2" needs no corpus, and a pipeline that retrieves unconditionally wastes latency and pollutes the context.
Agentic RAG lets a loop plan several retrievals, decomposing a compound question into sub-questions. Strictly more powerful and much harder to bound — without a step limit it can run for a long time on a query it cannot satisfy.
CRAG is the cheapest of the four and needs no training, which is why it is usually the first one to add. The others are what you reach for once you have measured that grading alone is not enough.
Checking the retrieval before trusting it
A plain RAG pipeline retrieves, then generates. It has no step that asks whether the retrieved documents are any good, so when retrieval fails the model still produces an answer — fluently, with citations, from irrelevant context.
Corrective RAG adds an evaluation step between the two:
query → retrieve → grade the documents → act on the grade → generate
The grader assesses each retrieved chunk for relevance to the question, and the pipeline branches on the result:
| Grade | Action |
|---|---|
| All documents relevant | Generate normally |
| Some relevant | Discard the rest, generate from what remains |
| None relevant | Do something else — rewrite, search elsewhere, or refuse |
That third row is where the value is. A plain pipeline's worst case is a confident answer from unrelated documents; a corrective pipeline's worst case is an honest refusal or a second attempt.
Grading the documents
The grader is a cheap model call per chunk, or one call over all of them:
Question: {question}
Document: {chunk}
Is this document relevant to answering the question? Reply with JSON:
{"relevant": true|false, "reason": "<one short sentence>"}Three design choices matter here:
Grade relevance, not correctness. The grader is deciding whether the chunk is about the question, which is a much easier and more reliable judgement than whether it answers it.
Use a small fast model. This runs per chunk on every query, so it is in the latency path. A small model does this adequately.
Batch the calls. Grading five chunks in one call is cheaper and faster than five calls, at some loss of independence.
The trade is explicit: an extra 200–600ms and a per-query cost, in exchange for knowing whether the context is usable before answering from it.
What to do when retrieval failed
The corrective branch is the interesting part, and there are four responses in increasing order of cost.
Refuse. "The available documents do not cover this." Honest, instant, and often the correct answer. Any corrective system should have this as its floor.
Rewrite the query and retry. The failure may be phrasing. Ask a model to reformulate and search again — one extra round trip, and it recovers a meaningful share of failures.
Broaden the search. Drop filters, raise k, add keyword search, try a different index.
Fall back to an external source. Web search, or another corpus. This is what the original Corrective RAG paper proposed for the no-relevant-documents case, and it is appropriate when the corpus is genuinely incomplete rather than badly searched.
A cap on retries is essential. Two attempts then refuse is a reasonable policy; without a limit, a hard query can loop.
What the grader costs, and what it buys
The sections above describe the three verdicts and what correction means. The question they leave open is whether the grader is worth its latency, and that is arithmetic -- so here it is, run over a set of queries where retrieval succeeds, partly succeeds and fails outright.
Things to try
- Select capital of Peru. Nothing clears the threshold, and the pipeline declines — the answer naive RAG cannot give.
- Raise the relevance threshold to 0.8 on a query that was working. More queries take the corrective path; grading is a precision/recall trade like any other classifier.
- Lower the ambiguous floor to 0.05. Weak evidence is now treated as partial rather than absent, which changes which path a borderline query takes.
What to remember
Corrective RAG grades retrieved documents before generating, and takes a different path when they are poor: rewrite the query, fall back to another source, or decline. Retrieval never fails loudly — it always returns its nearest k — so without a grader a query with no answer still produces a confident, sourced-looking one.
The family of self-checking pipelines
Corrective RAG is one of several patterns that add evaluation loops, and they check different things.
| Pattern | Checks | Acts by |
|---|---|---|
| Corrective RAG | Are the retrieved documents relevant? | Re-retrieving or falling back |
| Self-RAG | Is retrieval needed? Is the answer supported? | Retrieving on demand, critiquing output |
| Adaptive RAG | How hard is this question? | Routing to a cheaper or richer path |
| Agentic RAG | Anything, iteratively | A planning loop with tools |
Self-RAG adds two further checks: whether retrieval is needed at all (many questions do not need it), and whether the generated answer is supported by what was retrieved — regenerating if not.
Adaptive RAG classifies the query first. A simple factoid goes straight to single-pass retrieval; a complex multi-hop question goes to a decomposition path. This saves latency on the easy majority.
Agentic RAG is the general case: a loop where the model decides what to retrieve, evaluates what it got, and iterates. Most capable, least predictable, and hardest to bound in cost.
The common thread is replacing a fixed pipeline with a conditional one. The cost is latency and unpredictability; the gain is not answering confidently from nothing.
The cost, and when it is worth paying
| Plain RAG | Corrective RAG | |
|---|---|---|
| Model calls per query | 1 | 2–4 |
| Latency | Baseline | +30–100% |
| Answers from bad context | Yes | Rarely |
| Refuses appropriately | Only if prompted | By design |
| Predictable cost | Yes | No — depends on branching |
Worth paying when a wrong answer is expensive — regulated advice, medical or legal information, anything customer-facing where a confident error damages trust. Also worth paying when the corpus has known gaps, so refusal is a frequent correct answer.
Not worth paying when latency is tight, when retrieval quality is already high (measure recall@k first), or when the cost of a wrong answer is low. A well-tuned plain pipeline with a good refusal instruction covers a great deal of the same ground for a fraction of the latency.
That ordering matters: fix retrieval before adding a corrective loop. Grading documents is a way of detecting poor retrieval, not a substitute for improving it. If recall@k is 0.6, the grader will simply confirm that most retrievals are bad.
Implementation sketch
def corrective_rag(question, max_retries=1):
query = question
for attempt in range(max_retries + 1):
docs = retrieve(query, k=10)
keep = [d for d in docs if grade(question, d)["relevant"]]
if keep:
return generate(question, keep)
if attempt < max_retries:
query = rewrite(question) # try a different phrasing
else:
return refuse(question) # honest floorTwo things to log for every request: how many documents were graded relevant, and which branch was taken. Those two numbers turn the pipeline into a monitoring signal — a rising share of no-relevant-documents branches is an early warning that the corpus or the index has drifted, and it is visible before users complain.
Questions people ask
Does the grader need a large model? No. Judging topical relevance is a narrow task, and a small fast model is adequate and much cheaper in the latency path.
How much latency does it add? One batched grading call, typically 200–600ms, plus a full retry cycle when it triggers.
Is it better than reranking? They do different things. A reranker reorders; a grader decides whether to proceed at all. Use both — rerank, then grade the top few.
What if the grader is wrong? A false negative discards a good document; a false positive lets a bad one through. Validate the grader against human judgement on a sample, as with any model judge.
Should it fall back to web search? Only if the corpus is genuinely expected to be incomplete, and be explicit in the answer about which source was used.
Is this the same as agentic RAG? Corrective RAG is a fixed, bounded loop. Agentic RAG is an open-ended one with tools and planning.
Recap in one screen
- Corrective RAG grades the retrieved documents before generating, and branches on the result.
- Its floor is an honest refusal instead of a confident answer from irrelevant context.
- The corrective branch can rewrite the query, broaden the search, or fall back to another source — with a retry cap.
- It costs one extra model call plus occasional retries; worth it where a wrong answer is expensive.
- Fix retrieval first — grading detects poor retrieval, it does not improve it.