Recursive chunking
Split on the largest natural boundary that fits. Try paragraphs; if a piece is still over the limit, split it on sentences; then on words; then, as a last resort, mid-word. Every fallback loses a little more meaning, so the recursion only descends when it has to.
Overview
Why not just split every N characters
Fixed-size splitting is one line and cuts wherever it lands — mid-sentence, mid-word, mid-number. The retrieved chunk then starts halfway through a thought, and the generator has to answer from a fragment.
Recursive splitting keeps the size limit and spends it on the best boundary available. Chunks vary in length, which is fine: a chunk is a unit of meaning, not a unit of storage.
Parameters
Visualisation
—Readout
What to watch
- Each level is tried only when the level above left something too big.
- Fixed-size splitting cuts mid-sentence; this cuts at a boundary.
- The separator list is where domain knowledge goes.
Recursive chunking: A Practical Guide
What is recursive character chunking, and why is it the default?
The separator list is the whole configuration
The default is roughly ["\n\n", "\n", ". ", " ", ""] — paragraph, line, sentence, word, character. It descends only when a piece still exceeds the limit, so the last entry fires only on text with no whitespace at all.
Change the list to match the document type. Code wants ["\nclass ", "\ndef ", "\n\n"]; Markdown wants heading markers first. That is the cheapest large improvement available to a RAG pipeline, and it is usually left at the default.
Overlap, and what it costs
Chunks usually overlap by 10–20% so a sentence spanning a boundary appears whole in at least one of them. The cost is real: overlap inflates the index, and duplicated text means near-identical chunks compete in the results, crowding out genuinely different ones.
Recursive chunking is the right default and it is still structure-blind — it does not know a heading from a sentence. That is what the structure-aware and semantic variants address.
Choosing the size, and why there is no default
Chunk size is a trade between two failures. Too small and a chunk lacks the context to be understood alone — a sentence about "the second condition" without the first. Too large and the chunk covers several topics, so its embedding is an average of all of them and matches none precisely, while also burning context budget on material the query did not need.
The useful framing is that a chunk should be one retrievable idea. For dense prose that is often a paragraph, 200–500 tokens. For reference material with short entries it is much smaller. For code it is a function, whatever that costs in tokens.
Do not pick from a blog post. Build a small evaluation set of real questions, measure recall@k at three or four sizes, and take the winner. The difference between 200 and 800 tokens on a real corpus is routinely larger than the difference between two embedding models.
Overlap, and the deduplication it forces
Overlap exists because a boundary can fall mid-argument however carefully it is chosen. Repeating the last 10–20% of each chunk at the start of the next means a straddling sentence appears whole somewhere.
It costs more than index size. Overlapping chunks are near-duplicates, so a query matching the overlapped region retrieves both, and your top 5 is really a top 3 with two copies. That crowds out genuinely different evidence, and it is the reason maximal marginal relevance and other diversity-aware selection strategies exist.
If overlap is doing a lot of work for you, that is usually a signal the separators are wrong rather than that more overlap is needed. Fixing the boundaries is cheaper than paying for redundancy on every query.
Tuning the separator list per format
The default list — paragraph, line, sentence, word, character — assumes prose. Changing it for the document type is the single cheapest improvement available, and it is almost always left alone.
Code: lead with \nclass , \ndef , \n\n, so a function stays whole and a chunk is a unit someone could actually read.
Markdown: lead with heading markers, which turns the recursive splitter into a cheap approximation of structure-aware chunking.
Transcripts: split on speaker turns before sentences, so a chunk holds one person's contribution rather than half of two.
CSV and logs: split on lines and never below, because a half-row is meaningless.
Every one of these is a few characters of configuration against an embedding-model migration, and they routinely produce more improvement.
Splitting on the largest natural boundary that fits
Fixed-size chunking cuts at character 500 whatever is there — mid-sentence, mid-word, between a heading and its content. Recursive chunking fixes that with one idea: try a list of separators in order, from coarsest to finest, and use the largest unit that fits the size limit.
The default separator list, in order:
["\n\n", "\n", ". ", " ", ""]
Paragraph break, then line break, then sentence end, then word boundary, then individual characters as a last resort.
The algorithm: try to split the text on the first separator. If the resulting pieces are all within the size limit, done. If a piece is still too large, recurse into it with the next separator. The empty string at the end guarantees termination — a 600-character word will be cut, because it has to be.
That is the whole method, and it is the sensible default for almost every corpus. It costs nothing extra, respects natural boundaries the vast majority of the time, and degrades gracefully.
Why boundaries matter
A chunk cut mid-sentence damages retrieval in two specific ways.
The embedding is of a fragment. "the entitlement is 39 weeks provided that the employee has completed two years of" is a partial thought, and its vector is correspondingly muddled — it does not sit where either complete idea would sit.
The retrieved text is unusable. Even if it matches, passing a truncated sentence to the model means the answer is incomplete, or the model completes it from imagination.
Compare with a chunk that ends at a paragraph break: complete thoughts, a sharp embedding, and text that can be quoted.
| Split at | Chunk quality |
|---|---|
| Character count | Fragments, broken words |
| Word boundary | Complete words, broken sentences |
| Sentence end | Complete sentences, possibly mid-topic |
| Paragraph break | Complete thoughts — the target |
| Section heading | Complete topics — better still |
Recursive splitting reaches for the highest row it can.
Size and overlap
from langchain_text_splitters import RecursiveCharacterTextSplitter
splitter = RecursiveCharacterTextSplitter(
chunk_size=1000, # characters, not tokens - check which your tool uses
chunk_overlap=150, # ~15%
separators=["\n\n", "\n", ". ", " ", ""],
length_function=len,
)
chunks = splitter.split_text(document)Size is a compromise between context and precision. 200–500 tokens (roughly 800–2,000 characters) is the usual range. Too small and the chunk lacks context; too large and the embedding averages over several topics.
Overlap of 10–20% means consecutive chunks share their boundary text, so a sentence spanning a split appears whole in at least one chunk. The cost is duplicated storage and occasionally two near-identical results, which deduplication or MMR handles.
One detail that catches people: character count is not token count. chunk_size=1000 in characters is roughly 250 tokens for English prose and considerably fewer for code or non-Latin scripts. If the downstream constraint is a token budget, measure in tokens — most splitters accept a token-counting length function.
The separator list, and what overlap actually costs
Recursive splitting is entirely configured by its separator list and two numbers. This runs the algorithm on text with real structure, shows what each separator in the list is doing, and prices the overlap that everyone turns on without measuring.
Things to try
- Drop the chunk size to 60. More pieces fall through to sentence and then word level, and chunks start ending mid-sentence.
- Switch separators to fixed size. Every boundary lands wherever the character count ran out — watch the mid-sentence count climb.
- Raise the size to 260 with recursive separators. One chunk holds the whole passage, which retrieves as a single coarse unit.
What to remember
Recursive chunking splits on the largest natural boundary that fits, descending through a priority list of separators only when a piece is still oversized. The separator list is the whole configuration, and tailoring it to the document type is the cheapest large improvement available to a RAG pipeline.
Custom separators for real formats
The default list is tuned for prose. Other content wants different boundaries, and specifying them is where most of the available improvement lies.
Markdown:
["\n## ", "\n### ", "\n\n", "\n", ". ", " ", ""]
Headings first, so a section stays whole where possible.
Code:
["\nclass ", "\ndef ", "\n\n", "\n", " ", ""]
Class and function boundaries first. A half-function is not retrievable knowledge, and this ordering keeps definitions intact.
HTML: strip tags first, or use a splitter that understands the tag structure. Splitting raw HTML on newlines produces chunks full of markup.
Transcripts: split on speaker turns ("\nSpeaker: ") before paragraphs.
Legal text: section and clause numbering ("\nSection ", "\n(") carries the structure.
Most splitter libraries ship language-aware variants for common programming languages, which encode exactly this idea.
Where it sits among the strategies
| Strategy | Cost | Quality |
|---|---|---|
| Fixed size | Lowest | Poor — broken sentences |
| Recursive | Lowest | Good — the sensible default |
| Structure-aware | Low | Better where markup exists |
| Semantic | High — an embedding per sentence | Best on unstructured prose |
| Parent-child | Moderate | Best of precision and context |
Recursive chunking is the right starting point because it is nearly free and captures most of the benefit. The upgrades worth considering after it:
Structure-aware if the documents have real headings — use them, and prepend the heading path to each chunk so it carries context.
Parent-child if chunks need to be small for precision and large for context. Index small, return the parent section.
Semantic only if the documents are continuous prose with no usable structure, and only after measuring that chunking is genuinely the bottleneck.
Common mistakes
- Confusing characters with tokens, so chunks are four times smaller than intended.
- No overlap, so a sentence spanning a boundary is lost from both chunks.
- Too much overlap (50%+), which duplicates storage and fills results with near-identical chunks.
- Using prose separators on code, splitting functions in half.
- Splitting tables, which destroys them. Extract tables separately and keep each whole.
- Not storing the source and position, so citations cannot be produced later.
- Chunking before cleaning, so boilerplate and navigation text become chunks.
Questions people ask
What chunk size should I use? Start at 300–500 tokens with 15% overlap, then evaluate recall@k on your own question set.
Is recursive chunking better than semantic? Cheaper and frequently comparable on structured documents. Semantic wins on unstructured prose — measure rather than assume.
Should overlap be characters or sentences? Either. Sentence-based overlap is cleaner because it never cuts mid-thought.
How do I handle very long paragraphs? The recursion handles it — it falls through to sentence and then word boundaries.
Can chunk sizes vary within one index? Yes, and it is often right: a table or a function is one chunk whatever its length.
Does the separator order matter? Substantially. It encodes which boundaries you consider most important, and putting headings first is usually the largest single improvement.
Recap in one screen
- Try separators from coarsest to finest, taking the largest unit that fits — paragraphs, then lines, then sentences, then words.
- The empty separator at the end guarantees termination on pathological input.
- 200–500 tokens with 10–20% overlap covers most corpora.
- Customise the separator list for the format: headings for Markdown,
def/classfor code, speaker turns for transcripts. - Measure in tokens if the downstream budget is tokens, and store source and position for citations.