Structure-aware chunking
Split along the document's own structure. Headings, list items, table rows and code fences are boundaries the author already placed, and they are free. The second half matters as much: carry the heading path into each chunk, so a fragment still knows which section it came from.
Overview
Boundaries you do not have to guess
A Markdown heading, an HTML <section>, a PDF outline entry, a slide break: each is a statement by the author that the subject changes here. Semantic chunking spends an embedding per sentence to infer what the markup already says.
Some structures must not be split at all. A table split down the middle loses its header row and becomes unreadable; a code block cut in half is not code. These are kept whole even when they exceed the size limit, or summarised separately.
Parameters
Visualisation
—Readout
What to watch
- The boundaries were written by the author, not inferred.
- Each chunk carries its heading path, so it is self-describing.
- Splitting a table or a code block mid-way destroys it.
Structure-aware chunking: A Practical Guide
What is structure-aware chunking, and why does it beat a character splitter on real documents?
Carrying the heading path
The single most valuable thing this approach enables is context injection.
A chunk taken from deep inside a handbook might read: "The entitlement is 39 weeks, of which 6 are at full pay." Retrieved on its own, it does not say what kind of leave, or for whom.
Structure-aware chunking knows exactly where it came from, so the path can be prepended:
Employee Handbook > Leave > Parental Leave > Entitlement
The entitlement is 39 weeks, of which 6 are at full pay.
That transforms the chunk in two ways. Its embedding now contains the topic words, so a query about parental leave matches it. And the retrieved text is self-contained, so the model can answer without guessing.
This is cheap, mechanical, and one of the highest-value improvements available to a retrieval pipeline over naive chunking.
import re
def chunk_markdown(text, max_chars=1500):
chunks, path, buf = [], [], []
for line in text.split("\n"):
m = re.match(r"^(#{1,6})\s+(.*)", line)
if m:
if buf:
chunks.append((" > ".join(path), "\n".join(buf)))
buf = []
level = len(m.group(1))
path = path[: level - 1] + [m.group(2)]
else:
buf.append(line)
if sum(len(b) for b in buf) > max_chars:
chunks.append((" > ".join(path), "\n".join(buf)))
buf = []
if buf:
chunks.append((" > ".join(path), "\n".join(buf)))
return chunks
What it needs from you
A parser per format. Markdown is easy, HTML is manageable, PDF is genuinely hard — a PDF has no structure, only positioned glyphs, so headings must be inferred from font size and spacing. Most RAG quality problems on PDFs are really extraction problems.
Sections also vary wildly in length, so structure-aware chunking is usually combined with a recursive splitter: split on structure first, then recursively split any section still over the limit, keeping the heading path on every resulting piece.
Extraction is the hard part, not splitting
The splitting is easy once the structure is known. Knowing it is the problem, and it varies enormously by format.
Markdown and HTML carry structure explicitly; a parser gives you the tree.
PDF has no structure at all — only positioned glyphs. Headings must be inferred from font size, weight and spacing, multi-column layouts interleave text if read naively, and tables lose their rows. Most RAG quality complaints about PDFs are extraction failures rather than retrieval failures, and swapping the embedding model will not touch them.
Office documents and slides sit in between: real structure exists in the file format, and most extraction libraries flatten it away before you see it.
Budget for extraction seriously. It is unglamorous and it decides the ceiling on everything downstream.
Tables, code and the things that must not be split
Some structures are atomic. A table cut between its header row and its data leaves rows of unlabelled numbers; a code block cut in half is not code; a numbered list split mid-way loses the numbering that made it readable.
The usual handling is to keep these whole even when they exceed the size limit, and to accept the occasional oversized chunk. For a genuinely large table, the better answer is often to store a generated summary for retrieval and the full table for generation — the summary is what matches the query, the table is what answers it.
The same reasoning applies to images and diagrams: retrieve on a caption or description, return the artefact. Once you separate what is embedded from what is returned, you are already at the parent-document pattern.
What the heading path does to retrieval
Prepending Refund policy > Exceptions to a chunk is often described as a readability improvement, and its larger effect is on retrieval.
The section's vocabulary is now inside the embedded text. A query using the words of the section — "refund exceptions" — matches a chunk whose body never uses either word, because the body says "digital goods are non-refundable once downloaded". Without the path, that chunk is close to unfindable by its own section's name.
Store the path as metadata as well as prepending it. As metadata it supports filtering ("only search the exceptions"), it gives you a citation to display, and it lets you reconstruct the document's hierarchy for a parent-document lookup.
One caution: on a deep hierarchy the path can grow long enough to dominate a short chunk's embedding, which is the dilution problem described under context-aware chunking. Two or three levels is usually the useful limit.
Using the boundaries the author already wrote
Every structured document contains explicit markers of where topics change: headings, sections, list items, table boundaries, code blocks. Structure-aware chunking uses those instead of guessing.
The advantage over recursive splitting is that a heading is a semantic boundary asserted by the person who wrote the document. No similarity computation can be more authoritative than that, and it is free to use.
| Format | Structural markers |
|---|---|
| Markdown | #, ##, ###, lists, code fences |
| HTML | <h1>–<h6>, <section>, <article>, <table> |
| Outline/bookmarks, font-size hierarchy | |
| Word | Heading styles |
| Code | Classes, functions, imports |
| Notebooks | Cells |
| Slides | Slide boundaries |
The corresponding rule: split on the highest-level heading that produces chunks within your size budget, and fall back to finer boundaries inside sections that are too large.
Handling elements that must not be split
Structure also tells you what to keep whole, which fixed-size chunking cannot know.
Tables. Splitting a table mid-row destroys it — the header is in one chunk and the data in another, and neither is interpretable. Extract each table as its own chunk, and store a short text description alongside it for embedding, since a grid of numbers embeds poorly.
Code blocks. Keep whole. A half-function is not retrievable knowledge.
Lists. Keep short lists whole. A list of eight steps split after step three is worse than useless, because the retrieved fragment looks complete.
Definitions and glossary entries. One entry per chunk is usually right.
Figures and captions. Keep the caption with its reference.
The general rule: any element whose meaning depends on its own completeness gets its own chunk, whatever the size budget says. Variable chunk sizes are correct here, not a compromise.
The heading path, and what it does to retrieval
Structure-aware splitting uses boundaries the author already wrote. The splitting is the easy half; the two things that decide whether it is worth doing are what you carry along with each chunk, and what you refuse to split.
Things to try
- Compare structure first with fixed size at the same chunk size. The character splitter cuts through the heading; the structure-aware one does not.
- Lower the chunk size until a section still exceeds it. Structure gives boundaries, not size control, so a recursive splitter has to finish the job.
- Note the heading path on each chunk. A fragment retrieved alone still says which section it belongs to.
What to remember
Structure-aware chunking splits along the document's own markup — headings, list items, table rows — because those boundaries were placed by the author and cost nothing to find. Carrying the heading path into each chunk is the half people miss: it makes a fragment self-describing and puts the section's vocabulary into the embedded text.
Extraction is the hard part
The method depends on being able to see the structure, and for several common formats that is where the work is.
Markdown and HTML are easy — the structure is explicit and parseable.
PDF is the difficult one. A PDF is a description of marks on a page, not a document tree. Headings must be inferred from font size, weight and position; multi-column layouts are frequently extracted in the wrong reading order; and tables often come out as a jumble.
The practical advice is to check the extracted text before indexing thousands of documents. Extract ten, read them, and look for interleaved columns, missing headings and mangled tables. A layout-aware extractor is worth the effort where quality matters.
Scanned documents need OCR first, which loses structure and introduces character errors. Expect to fall back to semantic or recursive chunking for these.
Word documents carry heading styles, which is usable structure if the author applied them — and many do not, using bold text instead.
Where it sits among the strategies
| Strategy | Uses | Cost | Best on |
|---|---|---|---|
| Fixed size | Character count | Lowest | Nothing, really |
| Recursive | Separator hierarchy | Lowest | A safe default |
| Structure-aware | The document's own markup | Low | Structured documents |
| Semantic | Embedding similarity | High | Unstructured prose |
| Parent-child | Two granularities | Moderate | Long structured documents |
Structure-aware chunking is the right default where structure exists, and it composes with the others:
- Split on headings first, then recursively within any oversized section.
- Split on headings, then use parent-child so small precise chunks return their section.
- Fall back to semantic chunking for documents whose extraction produced no usable structure.
That layered arrangement handles a mixed corpus, which is what real corpora are.
Metadata worth capturing while you have it
The structural pass is the moment when this information is available, and capturing it later means re-indexing:
- Heading path — prepend to the text and store as a field.
- Heading level — useful for weighting or filtering.
- Page number and character offsets — needed for precise citations.
- Element type — paragraph, table, code, list. Enables filtering to tables when a question asks for figures.
- Document title, date, source and access control.
Citations are the practical payoff: "Handbook, section 4.2, page 17" is verifiable, and "one of the retrieved chunks" is not.
Questions people ask
What if the document has no headings? Fall back to recursive splitting, or semantic chunking if it is continuous prose.
How do I handle a very long section? Split it recursively inside, keeping the heading path on every resulting chunk.
Should chunk sizes be uniform? No. A table, a code block or a glossary entry should be one chunk whatever its length.
Is prepending the heading path really worth it? Yes — it improves both the embedding and the usability of the retrieved text, for almost no cost.
What about PDFs? Use a layout-aware extractor and inspect the output. This is where most structure-aware pipelines actually fail.
Can I combine this with semantic chunking? Yes, and it is a good arrangement: headings first, semantic splitting within oversized sections.
Recap in one screen
- Headings and sections are boundaries the author asserted — more authoritative than any inferred split.
- Prepend the heading path to each chunk: it sharpens the embedding and makes the retrieved text self-contained.
- Keep tables, code blocks and short lists whole, whatever the size budget says.
- Extraction is the hard part, and PDFs are where it usually goes wrong — inspect before bulk indexing.
- Capture heading path, page, element type and permissions during the structural pass, because adding them later means re-indexing.