How you split documents decides what retrieval can ever find
Retrieval quality is capped by chunking. A fact split across two chunks is a fact your system cannot retrieve, no matter how good the embedding model is.
| Strategy | Splits on | Best for |
|---|---|---|
| Fixed size | A token count | Uniform prose where structure does not matter |
| Recursive | Paragraph, then sentence, then word | The sensible default for mixed documents |
| Structural | Headings, sections, list items | Manuals, docs, anything with a real outline |
| Semantic | Embedding distance between sentences | Dense text with no markup; slow to build |
| Row / record | One record per chunk | Tables, CSVs, product catalogues |
| Parent-child | Small chunks that resolve to a bigger parent | When you must match precisely but answer with context |
| Content | Chunk | Overlap |
|---|---|---|
FAQ / short answers |
150–300 tokens | 0–20 |
Documentation |
400–800 tokens | 50–100 |
Long-form prose |
600–1000 tokens | 100–150 |
Code |
One function or class | Repeat the signature and imports |
Tables |
One row, header repeated | 0 |
Recursive splitting at 400–800 tokens with 50–100 of overlap handles most documents. Then measure recall on real questions — the right size is a property of your content, not a constant.
No. Overlap costs storage and returns near-duplicate chunks that crowd out genuinely different results. Use enough to keep a sentence from being cut in half, not more.