Chunking Strategies
Splitting documents into retrievable pieces: fixed-size, overlap, recursive, structure-aware and semantic chunking, and how to choose a chunk size.
What is it?
After loading, you have documents that may be dozens or hundreds of pages long. You cannot embed a whole book as one vector and expect search to work, and you do not want to paste a whole book into every prompt. Chunking splits each document into smaller pieces called chunks; each chunk is embedded, stored and retrieved on its own.
Why not embed whole documents? An embedding is a single vector that summarises the meaning of its text (see embeddings). A vector for a 50-page handbook is a blurry average of holidays, expenses, security and parking. A question about parking matches it weakly, and even if it is retrieved you would have to send all 50 pages to the model. Smaller chunks give sharper vectors and less irrelevant text in the prompt.
Why not tiny chunks, then? A single sentence like 'It is limited to 5 days.' has lost its context: what is limited? Chunks that are too small match queries but do not contain enough information to answer them. Chunk size is a trade-off between precision (small, focused chunks match the query sharply) and context (large chunks carry enough surrounding information to be understood).
The main strategies:
- Fixed-size chunking - cut every N characters, words or tokens. Simple and predictable, but it cuts through sentences and ideas.
- Overlap - each chunk repeats the last part of the previous one (for example 800 characters with 150 overlap). A fact that straddles a boundary then appears whole in at least one chunk. Overlap costs extra storage and embedding, and can produce near-duplicate results.
- Recursive chunking - try to split on the biggest natural boundary first (blank lines between paragraphs), and only if a piece is still too big, split it on smaller boundaries (single newlines, then sentences, then words). This keeps paragraphs and sentences intact whenever possible and is a strong default.
- Structure-aware chunking - use the document's own structure: Markdown headings, HTML sections, PDF pages, code functions, FAQ question-answer pairs. Each section becomes a chunk (split further if too long), and the heading path ('Handbook > Leave > Carry-over') is stored as metadata and often prepended to the chunk text so the chunk is self-explanatory.
- Semantic chunking - embed each sentence and start a new chunk where the meaning shifts (where similarity between neighbouring sentences drops). It follows topics instead of lengths, at the cost of an embedding call per sentence during ingestion.
Sizing rules of thumb. Sizes are usually measured in tokens (see tokens-and-context-windows); for English text one token is roughly four characters. Common starting points are a few hundred tokens per chunk (roughly 200-800) with 10-20% overlap. Q&A over policies and FAQs favours smaller chunks; summarising or reasoning over long arguments favours bigger ones. Embedding models also have a maximum input length; text beyond it is silently truncated, so chunks must stay under that limit (small local models such as all-MiniLM-L6-v2 were trained on short passages, so keep chunks for them short).
Every chunk carries metadata: the source document's metadata plus its own position (chunk index, page, section, character offsets). That is how you cite, filter, deduplicate and re-index later.
There is no universally best chunk size. Pick a sensible default (recursive, ~500 tokens, ~15% overlap, headings prepended), then measure retrieval quality on real questions and adjust (see evaluating-rag).
Explain like I'm 10
Chunking is like cutting a long film into scenes for a highlights search. Cut into single frames and nobody can tell what is happening in any of them. Keep the whole film as one clip and every search returns two hours of footage. Cutting at scene changes (structure-aware or semantic) works best, and letting each clip start a few seconds before the previous one ends (overlap) means no line of dialogue is lost at a cut.
Examples
Fixed-size chunks with and without overlap
const text = "Our refund policy is simple. Customers can request a refund within " +
"30 days of purchase. Refunds go back to the original payment method. " +
"Digital downloads are not refundable once opened. Contact support to start a refund.";
function chunkWords(text, size, overlap) {
if (overlap >= size) throw new Error("overlap must be smaller than size");
const words = text.split(/\s+/).filter(Boolean);
const chunks = [];
for (let start = 0; start < words.length; start += size - overlap) {
chunks.push({ index: chunks.length, startWord: start, text: words.slice(start, start + size).join(" ") });
if (start + size >= words.length) break; // last window reached the end
}
return chunks;
}
const fact = "within 30 days of purchase";
for (const overlap of [0, 4]) {
const chunks = chunkWords(text, 12, overlap);
console.log("size=12 words, overlap=" + overlap + " -> " + chunks.length + " chunks");
for (const c of chunks) console.log(" [" + c.index + "] " + c.text);
const whole = chunks.some((c) => c.text.includes(fact));
console.log(" fact '" + fact + "' intact in some chunk? " + whole + "\n");
}Without overlap the key fact is cut in half at a chunk boundary, so no single chunk contains '30 days of purchase' together with 'refund within'. With a 4-word overlap the fact survives intact in one chunk. Real systems measure in tokens or characters, but the mechanism is the same.
Recursive splitter that respects paragraphs and sentences
function recursiveSplit(text, maxLen, seps) {
seps = seps || ["\n\n", "\n", ". ", " "];
if (text.length <= maxLen) return [text];
if (seps.length === 0) { // last resort: hard cut
const out = [];
for (let i = 0; i < text.length; i += maxLen) out.push(text.slice(i, i + maxLen));
return out;
}
const [sep, ...rest] = seps;
const chunks = [];
let current = "";
for (const part of text.split(sep)) {
const candidate = current ? current + sep + part : part;
if (candidate.length <= maxLen) { current = candidate; continue; } // keep packing
if (current) chunks.push(current);
if (part.length > maxLen) { // this piece alone is too big: go finer
chunks.push(...recursiveSplit(part, maxLen, rest));
current = "";
} else {
current = part;
}
}
if (current) chunks.push(current);
return chunks;
}
const doc = "Leave policy\n\nEmployees get 25 days of annual leave. Up to 5 unused days carry over. " +
"Carry-over days expire on 31 March.\n\nSick leave\n\nSick leave is unlimited but needs a note after 3 days.\n\n" +
"Remote work\n\nYou may work remotely up to 3 days a week with approval.";
recursiveSplit(doc, 90).forEach((c, i) => console.log(i, "(" + c.length + " chars)", JSON.stringify(c)));Short paragraphs are packed together until the limit; a paragraph that is too long is split at sentence boundaries instead of mid-word. This is the idea behind the popular 'recursive character text splitter' found in many RAG libraries. Look closely at the output for two flaws real splitters handle: the '. ' separator is consumed, so a chunk can lose its final full stop (production splitters keep separators attached), and the heading 'Remote work' is packed at the end of the previous chunk, separated from its own text. The next example fixes the heading problem by chunking on structure.
Structure-aware chunks with the heading path prepended
const markdown = [
"# Employee Handbook",
"## Leave",
"Employees get 25 days of annual leave.",
"### Carry-over",
"Up to 5 unused days carry over. They expire on 31 March.",
"## Expenses",
"Submit receipts within 30 days. Meals are capped at 40 per day.",
].join("\n");
function chunkByHeadings(md, source) {
const chunks = [];
let path = [];
let body = [];
const flush = () => {
if (body.length === 0) return;
const section = path.join(" > ");
chunks.push({
text: section + "\n" + body.join(" "), // heading context travels with the text
metadata: { source: source, section: section, chunk: chunks.length },
});
body = [];
};
for (const line of md.split("\n")) {
const m = line.match(/^(#{1,6}) (.+)$/);
if (m) { flush(); path = path.slice(0, m[1].length - 1).concat(m[2]); }
else if (line.trim()) body.push(line.trim());
}
flush();
return chunks;
}
for (const c of chunkByHeadings(markdown, "handbook.md")) {
console.log(JSON.stringify(c.metadata));
console.log(" " + c.text.replace("\n", " | "));
}The chunk 'Up to 5 unused days carry over' is ambiguous alone; with 'Employee Handbook > Leave > Carry-over' prepended, both the embedding and the LLM know it is about annual leave. Long sections would additionally be passed through the recursive splitter, each piece keeping the same heading prefix.
Semantic chunking with sentence-transformers (Python)
# pip install sentence-transformers numpy
import re
import numpy as np
from sentence_transformers import SentenceTransformer
model = SentenceTransformer("all-MiniLM-L6-v2") # 384-dimensional vectors
def semantic_chunks(text: str, threshold: float = 0.45, max_sentences: int = 12):
sentences = [s.strip() for s in re.split(r"(?<=[.!?])\s+", text) if s.strip()]
if not sentences:
return []
vectors = model.encode(sentences, normalize_embeddings=True)
chunks, current = [], [sentences[0]]
for i in range(1, len(sentences)):
similarity = float(np.dot(vectors[i - 1], vectors[i])) # cosine, since normalised
topic_shift = similarity < threshold
if topic_shift or len(current) >= max_sentences:
chunks.append(" ".join(current))
current = []
current.append(sentences[i])
chunks.append(" ".join(current))
return chunks
text = ("Annual leave is 25 days. Unused leave can carry over up to 5 days. "
"The office kitchen is cleaned every Friday. Please label your food in the fridge.")
for c in semantic_chunks(text):
print("-", c)Adjacent sentences about the same topic have similar embeddings; a drop in similarity marks a topic change and starts a new chunk. The threshold must be tuned per corpus and embedding model, and a max size still applies so a long single-topic section does not become one giant chunk.
How it works
All chunkers share a loop: walk through the text, accumulate units (characters, tokens, words, sentences or sections) until a size limit or a boundary is reached, emit a chunk with metadata, then step back by the overlap and continue.
Measuring size. Character counts are fast and predictable; token counts match what the embedding model and LLM actually see. A practical approach is to chunk by characters with a conservative limit, or use the embedding model's own tokenizer to count tokens exactly.
Recursive splitting is a divide-and-conquer algorithm: split by the coarsest separator, greedily pack consecutive pieces up to the limit, and recurse with finer separators only on pieces that are still too big. The result respects natural boundaries whenever it can.
Overlap is applied when moving to the next chunk: the new chunk starts overlap units before the previous one ended. With chunk size S and overlap O, a document of length L produces roughly (L - O) / (S - O) chunks, so 20% overlap means about 25% more chunks to embed and store.
Contextual enrichment. Because chunks are retrieved in isolation, it helps to make each one self-contained: prepend the document title and heading path, or (a more advanced technique) have an LLM write a one-sentence context for each chunk before embedding it (see advanced-rag-techniques).
Parent-child chunking separates what you search from what you send: you embed small chunks for precise matching, but when one matches you send its larger parent section to the LLM (also covered in advanced-rag-techniques).
document (8 pages)
|====================================================|
fixed-size, no overlap:
|--c0--|--c1--|--c2--|--c3--|--c4--|--c5--|--c6--|
^ a fact cut in half at a boundary
fixed-size with overlap:
|--c0--|
|--c1--|
|--c2--| each chunk repeats the
|--c3--| tail of the previous one
structure-aware:
| # Intro | ## Leave | ### Carry-over | ## Expenses |
c0 c1 c2 c3
+ heading path stored in metadata / prependedWhy does it exist?
Embedding models produce one vector per input and work best on focused passages; LLM prompts have a finite context window and every token costs money and attention. Chunking is the compromise that makes large document collections searchable at a useful granularity: small enough to match precisely, big enough to be understood.
When to use it
Whenever documents are longer than a few paragraphs. Start with recursive chunking plus heading prefixes for prose, structure-aware chunking for Markdown, HTML, FAQs and code, and consider semantic chunking for long unstructured text where topics drift. Tune size and overlap with a retrieval evaluation set, not by intuition.
When not to use it
Do not chunk items that are already small and self-contained: a product description, a support ticket, an FAQ entry, a single row - index each as one unit. Do not split code in the middle of a function or a table in the middle of a row. If the entire corpus fits comfortably in the prompt, you may not need chunking (or retrieval) at all.
Common mistakes
Using one huge chunk per document, which produces blurry embeddings and stuffs the prompt with irrelevant text.
Using tiny chunks (single sentences) that match queries but lack the context needed to answer.
Setting overlap greater than or equal to chunk size, which causes an infinite loop or massive duplication.
Exceeding the embedding model's maximum input length, so the end of each chunk is silently truncated and never searchable.
Losing headings and titles, so a chunk like 'This is capped at 5 days' has no subject.
Not storing chunk position metadata (page, section, index), making citations and neighbour expansion impossible.
Choosing a chunk size once by gut feeling and never measuring retrieval quality.
Practice exercises
- Easy:
In the fixed-size demo, find the smallest overlap (in words) that keeps the fact intact for size 12. Then try size 8 and repeat.
- Easy:
Compute how many chunks a 100,000-character document produces with size 1,000 and overlap 0, 100 and 200 characters, using (L - O) / (S - O).
- Medium:
Modify recursiveSplit to add overlap: each chunk after the first should start with the last 20 characters of the previous chunk, cut at a word boundary.
- Medium:
Combine the two later demos: chunk by headings, then pass any section body longer than 80 characters through recursiveSplit, prefixing every piece with its heading path.
- Hard:
Build a code-aware chunker for JavaScript source: split at top-level 'function' declarations so each function is one chunk, with metadata {file, functionName, startLine}.
- Hard:
Take 20 real questions about a document you own. Chunk it at 200, 500 and 1,000 tokens, retrieve top-3 for each question with the Python embedding code, and record how often the chunk containing the answer appears. Which size wins?
Interview questions
How do you choose chunk size?
Start from a sensible default (a few hundred tokens with 10-20% overlap, boundaries at paragraphs/sections), stay under the embedding model's input limit, then measure: build a set of real questions with known answer locations and compare recall@k across sizes. Smaller chunks give precise matches for fact lookups; larger ones keep context for reasoning-heavy questions.
What problem does overlap solve, and what does it cost?
It prevents information that spans a chunk boundary from being split so that no chunk contains it whole. The cost is more chunks to embed and store (about S/(S-O) times as many) and near-duplicate results that can crowd the top-k, which you can mitigate by deduplicating or merging adjacent hits.
What is recursive chunking?
Splitting with a hierarchy of separators: try paragraph breaks first, greedily pack pieces up to the size limit, and only split oversized pieces further by line, sentence, then word. It keeps natural units intact while guaranteeing a maximum size.
Why prepend headings or titles to chunks?
Chunks are retrieved in isolation. A chunk like 'Up to 5 days carry over' does not say what it refers to; adding 'Handbook > Annual leave > Carry-over' makes the embedding match leave-related queries and lets the LLM interpret it correctly.
When would you use semantic chunking?
For long unstructured text where topic boundaries do not align with formatting, such as transcripts or essays. It embeds sentences and splits where adjacent-sentence similarity drops. It costs more at ingestion and needs a tuned threshold plus a size cap.
How should code and tables be chunked?
By their structure: code at function or class boundaries (often using a parser), tables by rows with headers repeated in each chunk, or stored in a database and queried rather than embedded. Splitting them mid-unit destroys meaning.