RAG · DOES THE ANSWER SURVIVE CHUNKING
RAG Chunker Visualizer
Chunk a document four ways and check whether each known answer still sits inside a single chunk, the condition for retrieving it whole.
Your experiment
Start with fixed 200-character chunks of the runbook and no overlap. Find the two split answers, then add 50 and 100 characters of overlap, switch to whole sentences and heading sections, try sizes 150, 250 and 300, and load the travel policy.
Every input recomputes the result immediately; there is no animation because nothing here unfolds over time. An input outside its allowed range is rejected with a message and the previous valid result stays on screen.
Chunks and the answer key
Computed data
Metrics
An answer counts as intact when its full text lies inside one chunk, which is what a retriever needs to return it whole; the four answers per sample document were chosen by hand and one of the runbook answers spans two sentences. Sizes are in characters; a sentence longer than the size becomes a chunk of its own. Overlap repeats the end of one chunk at the start of the next, so the stored text grows. Retrieval quality also depends on embeddings and ranking, which are not modelled.
Why chunk boundaries matter
A retriever returns chunks, not documents. If the sentence that answers a question is cut in two, neither half may rank highly and the model sees only part of the answer. This tool checks four hand-chosen answers per sample document and reports whether each lies entirely inside one chunk. With fixed 200-character chunks and no overlap the runbook gives SPLIT ANSWER: 2 of 4, because answer A1 crosses chunks 1 and 2 and A2 crosses chunks 2 and 3.
Overlap helps, at a price
Overlap repeats the end of each chunk at the start of the next. With 50 characters the single-sentence answer A1 is rescued and only A2 is still split, but there are 7 chunks instead of 5 and 1.3x the document is stored. At 100 characters it is 9 chunks and 1.81x, and A2, a 158-character answer spanning two sentences, is still split, because an answer longer than the overlap can always straddle a boundary.
Respect sentences and sections
The Chunking strategy decides where a chunk may end: fixed characters cut anywhere, whole words only between words, whole sentences only at sentence ends, and the heading strategy also never crosses a heading. Packing whole sentences up to the size keeps every runbook answer intact at 200 characters, with 7 chunks of 90 to 197 characters and no extra storage. Chunking within heading sections does the same and never mixes two sections. Both still fail when the size is too small for the two-sentence answer: at 150 characters, and for sentences also at 250, where the packing happens to start a chunk between its two sentences. Fixed chunks of 300 characters are lucky on this document: ALL 4 ANSWERS INTACT.
A second document
Load the travel policy. Fixed 200-character chunks split 1 of 4 answers, the hotel limits sentence. Whole words split the same sentence at 200, while whole sentences keep all 4 intact. At 100 characters the fixed strategy splits 2 of 4. The lesson carries over to your own documents: measure answer survival on questions you care about before choosing a chunk size, rather than trusting a default.
Reading the tool and its limits
The bars show chunk lengths, green when a chunk holds a whole answer and red when it holds part of a split one; the lanes list each answer's fate; the table gives every chunk's range; and the text panel prints the chunks and the answer key. Your own text has no answer key, so only the chunking is shown. Sizes are characters, not tokens, and embedding, ranking and reranking are not modelled.