Say you are building your first retrieval-augmented generation pipeline. You have a folder of PDFs, a vector database, and an embedding model. You chunk the documents into 500-character pieces, embed them, wire up a similarity search, and ask a question. The answer comes back subtly wrong: it cites a policy that belongs to a different section, or it stitches together two unrelated paragraphs and presents the blend as fact. The model is not broken. The chunking is.

Chunking is the step where you decide how to cut a long document into the units that get embedded and retrieved. Every downstream property of a RAG system depends on it. Retrieval precision, context length, token cost, and answer faithfulness all trace back to where you put the boundaries. A naive splitter that cuts at a fixed character count will slice a table in half, separate a heading from the paragraph it labels, and place a pronoun in one chunk while its referent sits in another. When the retriever later pulls that chunk, it hands the model a fragment that is internally coherent but externally incomplete.

This tutorial walks through the chunking step as a sequence of checkpoints. At each checkpoint you will see a symptom you are likely observing, the cause behind it, and the specific fix to apply. The code is copy-pasteable, the trade-offs are stated, and the failure modes are called out where they matter.


Checkpoint 1: Retrieved chunks read as fragments, not answers

You ask a question and the retrieved context starts mid-sentence, references “the above configuration,” or ends abruptly before the resolution. The model then either asks a clarifying question or invents a resolution.

Cause: You are splitting on a fixed character or token count with no awareness of document structure. This is the default behavior of most naive pipelines: text[i:i+500] in a loop. It is the fastest thing to write and the worst thing to ship, because a fixed window has no way to know that a sentence boundary, a heading, or a list item is a more meaningful cut point than character 500.

Fix: Move to a recursive splitter that tries a prioritized list of separators before falling back to a hard cut. The idea is to attempt to split on paragraph breaks first, then line breaks, then sentence punctuation, then spaces, and only cut mid-word as a last resort. LangChain’s RecursiveCharacterTextSplitter implements exactly this pattern, and the logic is simple enough to reproduce.

from langchain_text_splitters import RecursiveCharacterTextSplitter

splitter = RecursiveCharacterTextSplitter(
    chunk_size=800,
    chunk_overlap=120,
    separators=["\n\n", "\n", ". ", " ", ""],
    length_function=len,
)

docs = splitter.create_documents([raw_text])
for i, d in enumerate(docs[:3]):
    print(f"--- chunk {i} ({len(d.page_content)} chars) ---")
    print(d.page_content[:200])

Two parameters matter here. chunk_size is the target maximum length in characters (or tokens, if you pass a token-counting length_function). chunk_overlap repeats the tail of one chunk at the head of the next so that a fact spanning a boundary appears in full in at least one chunk. The separators list is ordered from coarsest to finest; the splitter descends the list only when a chunk still exceeds the size target.

Verify the result: print the length distribution of your chunks. If most chunks are far below chunk_size, your separators are too aggressive and you are producing tiny fragments. If many are exactly at chunk_size, you are falling through to the hard cut and losing structure.

lengths = [len(d.page_content) for d in docs]
print(f"count={len(lengths)} min={min(lengths)} max={max(lengths)} mean={sum(lengths)/len(lengths):.0f}")

A healthy recursive split typically shows a mean somewhere between 60 and 90 percent of chunk_size, with a long tail near the target for the documents that have no natural breaks.


Checkpoint 2: Tables, code, and lists come back mangled

A budget table retrieves as a column of numbers with no header. A code example retrieves as the first half of a function. A numbered procedure retrieves as steps 1 through 3 with the conclusion missing.

Cause: The recursive splitter treats every character equally. It has no model of a table or a fenced code block, so it will happily cut inside one. Prose survives this reasonably well because sentences are short and redundant; structured content does not, because a table header is the only thing that tells you what a column means.

Fix: Pre-process the document so that structured blocks are treated as atomic. Split on structure first, then recurse within each prose block. A practical approach is to split the raw markdown on blank lines into block-level segments, route non-prose blocks (tables, fenced code) through untouched, and only run the recursive splitter on paragraphs.

import re

FENCE = re.compile(r"```.*?```", re.DOTALL)
TABLE = re.compile(r"(?:^\|.*\|\s*$\n?)+", re.MULTILINE)

def segment_blocks(text: str) -> list[dict]:
    """Split text into typed blocks: 'code', 'table', 'prose'."""
    blocks = []
    protected: list[tuple[int, int, str]] = []

    for m in FENCE.finditer(text):
        protected.append((m.start(), m.end(), "code"))
    for m in TABLE.finditer(text):
        protected.append((m.start(), m.end(), "table"))

    protected.sort()
    cursor = 0
    for start, end, kind in protected:
        if start > cursor:
            blocks.append({"kind": "prose", "text": text[cursor:start]})
        blocks.append({"kind": kind, "text": text[start:end]})
        cursor = end
    if cursor < len(text):
        blocks.append({"kind": "prose", "text": text[cursor:]})
    return blocks

Then apply the recursive splitter only to the prose blocks. Code and tables pass through whole, or, if they exceed your token budget, get split at their own internal boundaries (a table by row groups with the header repeated, a code block by top-level function). This one change resolves the majority of “the retrieved chunk makes no sense” complaints, because the units the retriever sees now correspond to units a human wrote deliberately.

When not to do this: if your corpus is plain prose with no structure to protect, the extra segmentation adds latency for no benefit. Measure whether structured blocks are common before adding the pre-pass.


Checkpoint 3: The right document is retrieved, but the wrong part of it

Your vector store contains a 40-page product manual. The user asks about the return policy. Retrieval surfaces the manual, and the chunk you get back is from the introduction. The answer is somewhere in section 7, but the embedding of the introduction is close enough to the query vector that it wins.

Cause: You embedded chunks without any metadata about where they came from, and the query matched the document-level topic rather than the section-level one. The introduction of a manual is full of generic product language that overlaps with almost any question about the product.

Fix: Attach metadata to every chunk at split time, and use it to filter or rerank. At minimum, capture the source file, the section heading path, and a chunk index. Then either pre-filter the vector search by section or feed the heading text into the embedded content so the chunk carries its own context.

def attach_metadata(docs, source_path, heading_stack):
    """heading_stack is e.g. ['Product Manual', '7. Returns']."""
    for i, d in enumerate(docs):
        d.metadata = {
            "source": source_path,
            "heading_path": " > ".join(heading_stack),
            "chunk_index": i,
        }
        # Prepend the heading path so the embedding sees the localized topic.
        d.page_content = f"[{d.metadata['heading_path']}]\n{d.page_content}"
    return docs

Prefixing the heading path to the chunk text is a small, cheap trick that has an outsized effect on retrieval relevance. The embedding now encodes “7. Returns” alongside the policy language, so a query about returns is far more likely to surface the correct section than the generic introduction. The trade-off is a handful of extra tokens per chunk, which raises embedding cost slightly and consumes a little context budget at generation time. For most corpora that is a good trade.

Failure mode to watch: if headings are auto-generated or inconsistent (a common problem with converted PDFs), the prefix can add noise. Check that your heading extraction is producing real section titles before committing to this pattern.


Checkpoint 4: Answers are correct but the citations point at the wrong place

The generated answer is right, but the cited chunk does not contain the sentence the model used. The chunk boundary fell between the sentence and its source, and overlap was too small to cover the gap.

Cause: chunk_overlap is set to zero, or to a value smaller than the typical sentence length. When a fact straddles a boundary, the half in chunk A is retrieved and the half in chunk B is not, so the model answers from an incomplete fragment and the citation reflects whichever half happened to match.

Fix: Set overlap to roughly 10 to 20 percent of chunk_size, and make it large enough to cover at least one complete sentence. For an 800-character target, 120 to 160 characters of overlap is a reasonable starting point. Note that overlap is not free: it multiplies your embedding count and storage by the overlap ratio, so doubling overlap from 10 to 20 percent roughly doubles the marginal cost of that step.

SymptomLikely causeFixCost to add
Chunks read as fragmentsFixed-length split with no separatorsRecursive splitter with prioritized separatorsNone, drop-in
Tables and code mangledSplitter has no block modelSegment blocks, split prose onlyOne pre-pass function
Wrong section retrievedNo location metadataPrefix heading path, filter by metadataFew extra tokens per chunk
Citations miss the sourceOverlap too small or zeroOverlap 10-20% of chunk sizeProportional storage increase
Too many near-duplicate hitsOverlap too largeReduce overlap, add dedupe at retrievalMinor rerank step

Checkpoint 5: Every query returns six near-identical chunks

The top results are all variants of the same passage, shifted by a few characters. The model receives redundant context and the actual answer, sitting in a different chunk just below the cutoff, never makes it in.

Cause: Overlap is working against you. When overlap is large relative to chunk size, adjacent chunks share so much text that their embeddings land almost on top of each other. The retriever then returns the same content repeatedly and wastes the top-k budget on duplicates.

Fix: Reduce overlap and add a deduplication step after retrieval. A simple approach is to keep the highest-scoring chunk from any group whose text similarity exceeds a threshold (for example, 0.9 Jaccard on token sets), then fill the rest of the top-k from the next distinct candidates.

from itertools import islice

def dedupe(results, threshold=0.9, k=5):
    """results: list of (text, score), pre-sorted by score desc."""
    kept = []
    for text, score in results:
        tokens = set(text.split())
        if any(
            len(tokens & set(prev.split())) / len(tokens | set(prev.split())) > threshold
            for prev, _ in kept
        ):
            continue
        kept.append((text, score))
        if len(kept) >= k:
            break
    return kept

This is the same overlap-versus-redundancy tension you saw in Checkpoint 4, viewed from the opposite side. There is no single overlap value that is optimal for every corpus; the honest answer is that you tune it, and the right value depends on how self-contained your sentences are. Dense technical writing tolerates smaller overlap than discursive prose.


Checkpoint 6: Semantic chunking sounds better but performs worse

You replace the recursive splitter with a semantic splitter that embeds every sentence, measures adjacent-sentence distance, and cuts where the topic shifts. Retrieval quality drops.

Cause: Semantic chunking optimizes for topical coherence within a chunk, but retrieval cares about matching a query to a chunk. If your semantic splits produce very small, tightly focused chunks, each one covers a narrow slice of the topic and a query that spans two slices fails to match either. The method also adds an embedding pass over the whole corpus before the first chunk is even produced, which raises ingestion cost substantially.

Fix: Reserve semantic chunking for corpora where topic drift within a document is severe and prose is the dominant content type. For mixed corpora, use recursive splitting with block protection and heading prefixes; it is cheaper, more predictable, and easier to debug. If you do use semantic chunking, set a floor and ceiling on chunk size so that it cannot produce one-sentence chunks, and combine it with the metadata prefix trick from Checkpoint 3.

When not to use it at all: any corpus with heavy structure (tables, code, forms) or any pipeline where ingestion latency is on the critical path.


The verification loop

Chunking is not a set-once parameter. The working method is a short loop you run per corpus:

  1. Split with your current parameters and print the chunk length distribution.
  2. Build a set of 20 to 30 representative queries with known correct source passages.
  3. Measure recall at k (does the correct passage appear in the top k?) and mean reciprocal rank.
  4. Inspect the actual retrieved chunks for the queries that failed, and look for a boundary that cut the answer in half.
  5. Adjust one parameter, rerun, and compare. Change overlap before chunk size; change chunk size before switching splitters.

The reason to change one thing at a time is that chunk size, overlap, separator list, and metadata prefix all interact. Moving two at once makes attribution impossible, and the failure will reappear later without an obvious cause.

A useful diagnostic is to dump the five worst-performing queries alongside their retrieved chunks and read them as a human. In practice, most chunking defects are visible on sight: a table without a header, a paragraph that starts with “Therefore” and has no antecedent, a heading stranded alone. The metrics confirm what the eye already suspects. Start there, fix the obvious cut, and rerun the loop. The first pass usually moves the needle more than any subsequent tuning.