Semantic chunking
Put the boundary where the topic changes, not where the character count runs out. Embed each sentence, measure the similarity between consecutive ones, and cut at the troughs. Chunks come out semantically coherent and wildly variable in length.
Overview
How the boundary is chosen
Split into sentences, embed each one, and compute the similarity between each consecutive pair. Where the text stays on topic the similarity is high; where the subject changes it dips. Cut at the dips.
The threshold should be relative. A fixed number like 0.7 means something different in every corpus, so implementations use a percentile of the similarities observed in this document — the 5th percentile of gaps, say — which adapts automatically to how varied the writing is.
Parameters
Visualisation
—Readout
What to watch
- The boundary lands at the similarity trough, not at a size limit.
- Chunk lengths become uneven, which is the point.
- The threshold is a percentile of this document, not a global constant.
Semantic chunking: A Practical Guide
What is semantic chunking, and when is it worth the extra cost?
What it costs
The cost is real and worth stating plainly: one embedding call per sentence at indexing time.
A 100,000-document corpus averaging 200 sentences each is 20 million embedding calls, against perhaps 400,000 for fixed chunking. That is a fifty-fold increase in indexing cost.
| Strategy | Indexing cost | Quality |
|---|---|---|
| Fixed size | Lowest | Poor — cuts mid-sentence |
| Recursive | Lowest | Good |
| Structure-aware | Low | Better where structure exists |
| Semantic | High | Best on unstructured prose |
Whether that trade is worth it depends on one thing: does the document have usable structure? If it has headings and sections, structure-aware chunking gets most of the benefit for almost none of the cost, because the author already marked where the topics change. Semantic chunking earns its cost on continuous unstructured prose — transcripts, scanned text, articles without headings.
When to reach for it
Worth it for long unstructured prose — transcripts, interviews, reports without headings — where no formatting signal exists and a fixed-size split reliably cuts mid-argument.
Not worth it when the document already carries structure. If there are headings, they are an explicit, author-provided, free topic boundary, and structure-aware chunking will beat semantic chunking at a fraction of the cost. Measure before adopting: on a retrieval evaluation set the gain over a well-tuned recursive splitter is often small.
Making the similarity signal usable
Raw sentence-to-sentence similarity is noisy. Short sentences — "Yes.", "See above.", a heading fragment — have unstable embeddings, so the gap sequence contains dips that are artefacts rather than topic changes.
Two standard mitigations. Buffering: embed each sentence together with its neighbours, so a short sentence inherits context and its vector stops swinging. Smoothing: take a rolling mean over the gap sequence before looking for troughs, so a single anomalous sentence cannot open a boundary on its own.
Both trade sensitivity for stability, and both are worth having. Without them the technique looks impressive on a clean essay and falls apart on a support transcript, which is exactly the sort of unstructured text it was supposed to be good at.
Is it worth it? Measure before adopting
Semantic chunking is the most cited of the strategies and the hardest to justify on evidence. It costs an embedding call per sentence at index time, produces chunks of unpredictable size that complicate context budgeting, and on published comparisons its advantage over a well-tuned recursive splitter is often small or absent.
Where it does earn its cost: long unstructured prose with no formatting signal — interview transcripts, meeting notes, scanned reports — and corpora where topics shift within a paragraph, so paragraph boundaries mislead.
The honest order of operations is to tune the separators and the size first, measure, and only then try semantic chunking against that baseline. It is a real technique that is frequently adopted before the cheap options have been exhausted.
How it relates to the other three strategies
The four strategies answer genuinely different questions, which is why they compose rather than compete.
Recursive asks where the safest boundary is given a size budget. It is the default and the baseline.
Structure-aware asks what boundaries the author already marked. Free when the format carries them, and strictly better than guessing when it does.
Semantic asks where the meaning changes. It is the only one that can find a boundary the author did not mark, which is why it is for unstructured prose specifically.
Context-aware asks what the chunk lost by being cut, and is orthogonal to all three — you can enrich chunks produced by any splitter.
A sensible production pipeline is usually structure first, recursive to enforce the size limit, and context enrichment on the result. Semantic chunking enters when the documents have no structure to exploit.
Splitting where the meaning changes
Fixed-size chunking cuts every 500 characters regardless of what is there. Recursive chunking respects paragraph and sentence boundaries. Neither knows whether a boundary is a topic boundary.
Semantic chunking finds those boundaries by measuring meaning directly:
- Split the document into sentences.
- Embed each sentence.
- Compute the similarity between each consecutive pair.
- Split where the similarity drops below a threshold.
Where two consecutive sentences are semantically close, they belong together. Where the similarity falls sharply, the topic has changed and that is where the chunk should end.
sim(sᵢ, sᵢ₊₁) high → same chunk sim low → boundary
The result is chunks that are each about one thing, at whatever length that takes — which is what an embedding wants, since a vector representing one topic is sharp and one representing three is a blurred average.
The threshold, and how to set it
An absolute similarity threshold does not transfer between documents or models, because typical similarity levels vary enormously.
The robust approach is a percentile of the distances within this document: compute all consecutive-sentence distances, and split at the ones above the 90th or 95th percentile.
| Percentile | Effect |
|---|---|
| 80th | Many boundaries, small chunks |
| 90th | A reasonable default |
| 95th | Fewer boundaries, larger chunks |
Two refinements that materially improve results:
Use a buffer. Compare a small window of sentences either side of the candidate boundary rather than single sentences. A single short sentence — "This is important." — produces a spurious dip on its own.
Enforce size limits. Semantic boundaries can produce a 4,000-token chunk or a 30-token one. Cap the maximum and merge below a minimum, or the downstream context budget becomes unpredictable.
import numpy as np
sents = split_sentences(text)
vecs = model.encode(sents, normalize_embeddings=True)
sims = [float(vecs[i] @ vecs[i+1]) for i in range(len(vecs)-1)]
threshold = np.percentile(sims, 10) # lowest 10% of similarities
boundaries = [i+1 for i, s in enumerate(sims) if s < threshold]
The similarity curve, and where it dips
Semantic chunking splits where consecutive sentences stop resembling each other. That description leaves the important question open -- how big a dip counts -- so here is the actual signal from a passage, and the boundaries each threshold produces.
Things to try
- Raise the split percentile to 70. It splits almost everywhere, producing chunks too small to carry an idea.
- Drop it to 10. Only the single largest topic change survives as a boundary.
- Note that the threshold is derived from this document's own gaps — a fixed number would behave differently on every corpus.
What to remember
Semantic chunking embeds each sentence and cuts where the similarity between neighbours drops, so boundaries land at topic changes rather than at character counts. It costs one embedding per sentence at index time, and it is usually not worth it on documents that already carry headings.
Where it helps and where it does not
It helps on:
- Meeting and interview transcripts, where topics shift with no markup.
- Long-form articles and books without useful headings.
- OCR output, where structure was lost in extraction.
- Documents whose headings are decorative rather than topical.
It does not help on:
- Well-structured technical documentation. The headings are better boundaries than any similarity measure, and free.
- Short documents already about one thing — FAQ entries, product descriptions, tickets.
- Highly uniform text, where consecutive-sentence similarity is flat and no clear boundaries exist.
- Code, where structural boundaries (functions, classes) are what matter.
There is a further honest caveat: the measured benefit is often smaller than expected. Published comparisons on well-structured corpora frequently find semantic chunking roughly level with recursive chunking. It is worth evaluating rather than adopting on principle.
Variants
Statistical / gradient-based. Instead of a percentile threshold, look for the largest local drops in similarity — the points of steepest change — which adapts better to documents with varying similarity levels.
Double-pass. Chunk semantically, then merge adjacent chunks whose embeddings are similar. Cleans up over-splitting.
LLM-based (agentic) chunking. Ask a model to propose boundaries, or to group sentences into topics. Highest quality on difficult text, and by far the most expensive — usually reserved for small high-value corpora.
Late chunking. A recent alternative worth knowing: embed the whole document with a long-context embedding model, then pool the token embeddings into chunk vectors. Each chunk's vector is informed by the entire document, which addresses the missing-context problem without needing per-sentence calls.
Practical guidance
Start with recursive or structure-aware chunking and a proper evaluation set. Measure recall@k. Then, if retrieval is the bottleneck and the documents lack structure, try semantic chunking and measure the difference.
That ordering matters because chunking is one of several levers, and it is not usually the largest. Hybrid search and reranking typically give bigger gains for less effort, and they are worth exhausting first.
Whatever strategy you choose, two things help regardless: prepend the heading path to each chunk so it carries context, and consider parent-child retrieval so small precise chunks can return larger context.
Questions people ask
Is semantic chunking worth the cost? On unstructured prose, often. On documents with real headings, usually not — measure it.
What embedding model for the sentence comparisons? A small fast one is adequate; it is comparing adjacent sentences, not doing retrieval.
How do I stop chunks becoming enormous? Enforce a maximum size and split at the weakest internal boundary when exceeded.
Does it replace overlap? Largely — boundaries at genuine topic shifts lose less than arbitrary ones. A small overlap is still cheap insurance.
Can I combine it with structure? Yes, and this is often the best answer: split on headings first, then semantically within any oversized section.
Is late chunking better? Promising, and it needs a long-context embedding model. Worth evaluating if one is available.
Recap in one screen
- Embed sentences, measure consecutive similarity, and split where it drops — boundaries at real topic changes.
- Use a percentile of the document's own distances rather than an absolute threshold.
- Add a comparison buffer and enforce minimum and maximum sizes.
- It costs one embedding call per sentence, which is a large indexing overhead.
- Worth it on unstructured prose; structure-aware chunking usually wins where headings exist.