Context-aware chunking

Every chunking strategy answers "where do I cut?". This one answers a different question: what did this chunk lose by being cut? A fragment full of "it" and "the above" is unretrievable and unusable, so context is added back — a title, a heading path, a generated summary, resolved pronouns.

Overview

The problem is coreference, not boundaries

Prose is written to be read in order, so it leans on everything before it: "it", "this policy", "as described above", "the latter". Cut one paragraph out and those references dangle.

Two consequences. The generator cannot answer from the chunk because it does not know what "it" is. And retrieval never surfaces the chunk in the first place, because the embedding is of text that does not contain the words a user would search for.

Parameters

Visualisation

—

Readout

What to watch

  • The boundary is unchanged; what is stored changes.
  • Pronouns and back-references are what make a raw chunk unusable.
  • The enrichment goes into the embedded text, so retrieval improves too.

Context-aware chunking: A Practical Guide

What is context-aware chunking, and what problem does it solve that the other strategies do not?

What gets added

Cheap and free: document title and heading path, prepended. Deterministic, no model call, and usually the largest single improvement.

Neighbour context: the previous and next chunk, or a window around the chunk, stored for generation but not embedded — the parent document retriever is exactly this idea, retrieving small and returning large.

Generated context: an LLM writes a sentence situating the chunk in the document, prepended before embedding. Anthropic's "contextual retrieval" is this, and it works well. It costs one model call per chunk at index time, which prompt caching over the shared document makes affordable.

What it costs, and what to watch

Index time and money, both once. The subtler cost is dilution: prepending 200 tokens of context to a 300-token chunk means the embedding is largely about the context, so chunks from the same section start looking alike and the retriever loses its ability to distinguish them.

Keep the enrichment short relative to the chunk, and measure on a retrieval evaluation set rather than assuming. This is the strategy with the best evidence behind it and it is not free of trade-offs.

Contextual retrieval, and what it costs

The strongest published form of this is to have an LLM write a short sentence situating each chunk inside its document — what it is about, what it follows — and prepend that before embedding. Reported reductions in retrieval failure are large, and the technique is simple enough to implement in an afternoon.

The cost is one model call per chunk at index time, which sounds prohibitive on a large corpus and mostly is not: the document is the same for every chunk in it, so prompt caching over that shared prefix makes the marginal call cheap. It is the clearest practical example of two of these techniques composing.

It is still an index-time cost paid on every re-index, so it belongs in your pipeline design rather than being bolted on: changing the chunker means regenerating every context sentence.

Embedded text and returned text can differ

The enrichment does not have to be what the model reads. You can embed the enriched text — so the chunk is findable — while storing and returning the original, so the generator is not fed repetitive boilerplate.

Separating the two is the general form of a pattern that appears everywhere in retrieval: embed something optimised for matching, return something optimised for reading. Parent-document retrieval embeds a small chunk and returns its parent. Summary indexing embeds a summary and returns the document. Contextual retrieval embeds chunk-plus-context and can return either.

Once you see the pattern, the design question stops being "how should I chunk?" and becomes two questions with different answers: what should be matched against, and what should be read?

Measuring whether the enrichment helped

Enrichment is easy to add and easy to overdo, and the only way to tell which you have done is to measure both halves.

Did findability improve? Recall@k on a fixed query set, before and after. This is the number the technique is supposed to move, and it usually does.

Did distinguishability degrade? The failure mode is chunks from one section becoming interchangeable. A cheap proxy is the mean pairwise similarity between chunks that share a parent: if it climbs sharply after enrichment, the shared context is dominating and the retriever is losing its ability to pick between them.

Watching only the first will lead you to enrich more and more, because recall keeps improving right up until the point where the top-k fills with near-identical chunks from the same section — at which point recall still looks fine and answers get worse.

The missing-context problem

A chunk is retrieved and passed to the model on its own. Frequently it cannot be understood on its own:

"This must be requested at least 8 weeks in advance, and the form requires a manager's signature."

What must? Requested from whom? The chunk before it said "Parental leave", and that chunk is not in the context.

This is the central weakness of chunking, and it has nothing to do with where the boundaries fall — even a perfectly-placed boundary produces chunks that reference their surroundings. Documents are written to be read in order; chunks are retrieved out of order.

Context-aware chunking is the family of techniques that give each chunk enough surrounding information to stand alone.

Four ways to add the missing context

TechniqueWhat is addedCost
Heading path prefixThe document's section hierarchyFree
OverlapThe neighbouring boundary textStorage
Parent-childRetrieve small, return the parent sectionStorage, complexity
Contextual summaryA generated sentence describing the chunk's placeOne model call per chunk

The heading path is the cheapest and often the most effective. Prepending "Handbook > Leave > Parental Leave > Applying" to that chunk resolves "this" immediately, improves the embedding, and costs nothing beyond capturing the structure at indexing time.

Overlap of 10–20% means a sentence spanning a boundary appears whole in at least one chunk. It handles the boundary case and not the "referenced three paragraphs earlier" case.

Parent-child indexes small chunks for precise matching and returns the larger enclosing section for context. It resolves references within the section by construction.

Contextual retrieval is the newest of the four: ask a model, at indexing time, to write a sentence or two situating each chunk within its document, and prepend that to the chunk before embedding.

Contextual retrieval, concretely

The prompt is short and the effect is substantial:

Here is a document:
<document>{full_document}</document>

Here is a chunk from it:
<chunk>{chunk_text}</chunk>

Write 1-2 sentences of context situating this chunk within the document,
so it can be understood on its own. Reply with the context only.

The generated context is prepended to the chunk text, and the combined text is embedded and stored.

Generated: "This passage is from the parental leave section of the employee handbook, describing the application process for eligible employees."
Chunk: "This must be requested at least 8 weeks in advance..."

Now the chunk is retrievable by a query about applying for parental leave, and interpretable when retrieved.

The reported effect on retrieval failure rates is large — substantial reductions when combined with hybrid search and reranking. The cost is one model call per chunk at indexing time, which for a large corpus is significant. Two mitigations: use a small fast model, and exploit prompt caching, since the full document is repeated across all of its chunks' prompts.

Testing whether a chunk can stand alone

The problem is coreference -- a chunk that says "it" or "the policy" is meaningless once separated from what it referred to. This measures that damage directly, applies the four repairs the article names, and shows which one is worth its cost.

example_01.pyNumPy
Output

Things to try

  1. Start with everything off. The raw chunk is about refunds, never says so, and cannot be retrieved.
  2. Turn on resolve pronouns. Similarity jumps, because the searchable nouns are now in the text.
  3. Turn on over-enrich. The chunk stays findable and stops being distinguishable from its sibling — the cost of adding too much shared context.

What to remember

Context-aware chunking changes what is stored rather than where the cut falls. A chunk full of pronouns and back-references is unusable by the generator and invisible to retrieval, so title, heading path and resolved references are added back. Keep the enrichment short relative to the chunk, or shared context dominates the embedding and chunks stop being distinguishable.

Choosing among them

The techniques are not alternatives so much as layers, and they have a clear cost ordering.

Always do: capture the structure and prepend the heading path. Free, and it addresses the most common case.

Usually do: 10–20% overlap. Cheap insurance against boundary loss.

Often worth it: parent-child retrieval, if chunks need to be small for precision. It resolves intra-section references without a model call.

Consider: contextual retrieval, when the corpus is high-value, the documents are long, and retrieval quality is the bottleneck. Measure the gain against the indexing cost.

Emerging: late chunking — embed the whole document with a long-context embedding model, then pool token embeddings into chunk vectors. Each chunk's vector is informed by the entire document without any per-chunk generation. Promising, and it requires a suitable embedding model.

Testing whether a chunk stands alone

There is a simple and effective test, and it is worth running on a sample before indexing a whole corpus.

Take twenty chunks at random and read each one cold. For each, ask:

  1. Can I tell what topic this is about?
  2. Do all the pronouns and references resolve?
  3. If someone asked a question this chunk answers, would this text answer it?

If the answer to any is no, the chunk needs more context. That is a five-minute exercise that predicts retrieval quality better than most metrics, and it directly identifies which of the four techniques above is needed.

Following it up quantitatively: measure recall@k with and without the context addition on a question set. The heading path alone frequently moves recall by several points.

Common mistakes

  • Adding context after embedding. The context must be present in the text that gets embedded, or it improves readability without improving retrieval.
  • Prepending so much context that it dominates the embedding. A 200-token context on a 100-token chunk means the vector is mostly about the context. Keep it to a sentence or two.
  • Not storing the original chunk separately, so citations quote generated context as if it were source text.
  • Generating context per chunk without prompt caching, paying for the full document repeatedly.
  • Assuming overlap solves reference resolution. It handles boundaries, not distant antecedents.

Questions people ask

Is the heading path enough? Often, for well-structured documents. It is always worth doing first because it is free.

How much does contextual retrieval cost? One model call per chunk. With a small model and prompt caching over the shared document, considerably less than it first appears.

Does the generated context risk hallucination? It can misdescribe a chunk. Keep it to situating rather than summarising content, and store the original text separately for citation.

Should context be embedded or only stored? Embedded — that is the point. Store the original separately for display.

Does this replace good chunking? No, it complements it. Bad boundaries plus added context is worse than good boundaries plus added context.

What is late chunking? Embedding the whole document first and pooling token vectors into chunk vectors, so each chunk's embedding already carries document context.

Recap in one screen

  • Chunks routinely reference material outside themselves, and retrieval delivers them out of order.
  • Prepend the heading path always — free, and it fixes the most common case.
  • Overlap handles boundary loss; parent-child retrieval handles intra-section references.
  • Contextual retrieval generates a situating sentence per chunk at indexing time, and measurably reduces retrieval failures.
  • Read twenty chunks cold and ask whether each stands alone — the cheapest diagnostic available.

Recall check

0 of 3

Say the answer out loud before you reveal it — recalling it is what makes it stick, and rereading it is not.

  1. Without scrolling back — what is the one-line takeaway from this module?

  2. What does this module say about “The problem is coreference, not boundaries”?

  3. What does this module say about “What gets added”?

Cheat sheet

Context-aware chunking

Every chunking strategy answers "where do I cut?". This one answers a different question: what did this chunk lose by being cut? A fragment full of "it" and "the above" is unretrievable and unusable, so context is added back — a title, a heading path, a generated summary, resolved pronouns.

GEN AI · vizlearn.in/gen_ai/context_aware_chunking.html

About the author

Ashish Jangra builds and maintains VizLearn. Every module here is written and the visualisation behind it hand-built, so the numbers in a readout come from the same code that draws the picture. Corrections are genuinely welcome and get priority over everything else — if a page states something wrong, or an animation misrepresents what the algorithm does, get in touch.