What is chunking in RAG? Chunk size, overlap, structure-aware splitting and the experiments that actually decide retrieval quality

Chunking splits documents into retrievable passages, and it decides RAG quality more than the model does. Fixed, recursive, semantic, structure-aware and parent-child chunking compared, how to choose chunk size and overlap from your own questions, how to chunk PDFs, tables and code, and the mistakes that make retrieval fail.

In a packed college auditorium
In a packed college auditorium

Key takeaways

  • Chunking decides what retrieval can ever return. A chunk must be able to answer a realistic question on its own.
  • Structure-aware splitting on headings, paragraphs, table rows and code blocks beats any fixed token count, and prepending the section heading to each chunk is the single highest-value trick.
  • There is no universal chunk size. Pick candidates, run twenty real questions, measure recall at k, and keep the winner.
  • Overlap is insurance for sentences that straddle boundaries, not a way to fix bad splitting. Keep it small.
  • Store rich metadata with every chunk: source, page, section, position. Citations and debugging depend on it.

What is chunking in RAG?

Chunking is splitting documents into passages, chunks, that are embedded and indexed individually so retrieval can return the piece of a document relevant to a question rather than the whole document. Every later stage of a RAG system, embedding, retrieval, reranking, generation, can only choose among the chunks you created, so chunking sets the ceiling on answer quality. A chunk that lost its context cannot be recovered by a better model.

Also asked as: what is chunking in rag · what is chunking in llm · chunking meaning in ai · chunking in nlp · what is text chunking · document chunking · chunking in generative ai · why chunking is important in rag · what is chunking and embedding

Most RAG quality problems I have debugged were chunking problems in disguise. The model looked like it was hallucinating; the retriever was handing it half a table, or a paragraph whose subject lived in the previous chunk. Barnett et al. list missing content and wrongly ranked passages among the seven failure points of RAG systems, and both usually trace back to how documents were split [3].

Before you tune prompts or switch models, print the chunks your retriever returned. Nine times out of ten, the bug is sitting right there in plain text. Pranjul Rathour, from building RAG.NextUpgrad

Why does chunking matter so much?

Three reasons. Embeddings represent one meaning per vector, so a chunk mixing several topics produces a blurred vector that matches nothing well. Retrieval returns whole chunks, so a chunk with the answer buried in unrelated text wastes the model's context and attention, and models attend poorly to the middle of long contexts [4]. And chunks are the unit of citation, so a chunk without its heading or page cannot be pointed at. Good chunks are precise, self-contained and traceable.

Also asked as: why chunking matters in rag · how chunking affects retrieval · chunking and embeddings · chunk quality rag · impact of chunk size on rag performance · chunking problems rag

What are the chunking strategies?

Fixed-size chunking splits every N tokens with an overlap and ignores structure; it is a baseline, not a strategy. Recursive splitting tries paragraph, then sentence, then word boundaries to hit a size. Structure-aware chunking splits on the document's own units, headings, paragraphs, list items, table rows, code blocks, and keeps them whole. Semantic chunking groups sentences by embedding similarity. Parent-child chunking indexes small chunks for precise retrieval but returns their larger parent for context. Most production systems use structure-aware splitting with a size cap and parent-child retrieval.

Also asked as: chunking strategies for rag · best chunking strategy for rag · chunking techniques in rag · fixed size chunking vs semantic chunking · recursive character text splitter · semantic chunking · parent child chunking · document chunking methods · chunking methods for llm

LangChain and LlamaIndex ship implementations of each [7][8]; Unstructured partitions PDFs, Word files and HTML into typed elements you can chunk structurally [9]. The survey by Gao et al. covers these and the research variants [2].

What is the best chunk size for RAG?

Whatever size lets a chunk answer a realistic user question on its own, and that differs by corpus. Policy documents and manuals often work at a few hundred tokens; API references at one function signature plus its description; legal clauses at the clause; support tickets at the ticket. Anyone quoting one number without asking about your documents is guessing. Pick three candidate sizes, index each, run twenty real questions, measure recall at k, and keep the winner.

Also asked as: best chunk size for rag · rag chunk size · chunk size and overlap · optimal chunk size · chunk size 512 vs 1024 · how to choose chunk size · chunk size experiments · small chunks vs large chunks rag · token chunk size rag

My own experiments on this, with the numbers for the corpora I tested, are written up separately [12]. The pattern that repeated: smaller chunks retrieve more precisely but need a parent or a heading to be understood; larger chunks are understood alone but retrieve less precisely. Parent-child gets both.

How much overlap should chunks have?

Small: enough to catch a sentence that straddles a boundary, roughly ten to fifteen percent of the chunk, and none at all when you split on structure that already respects sentence and paragraph ends. Large overlaps duplicate text, inflate the index, and make the reranker see the same passage twice. If you find yourself increasing overlap to fix retrieval, the real problem is where the boundaries fall, and the fix is structure-aware splitting.

Also asked as: chunk overlap rag · how much overlap for chunking · chunk overlap size · overlap in text splitting · why chunk overlap · sliding window chunking · chunk overlap best practice

Should I add context to each chunk?

Yes. Prepend the document title and the section heading path to every chunk, so a passage that says "the limit is 5 requests per second" is indexed and retrieved as "API Reference, Section 4.2 Rate limits: the limit is 5 requests per second". This single change did more for retrieval accuracy in RAG.NextUpgrad than any embedding model swap [13]. Contextual retrieval goes further by having a model write a short situating sentence for each chunk before embedding, at index-time cost [6].

Also asked as: contextual chunking · add metadata to chunks · heading in chunk · contextual retrieval anthropic · chunk context enrichment · chunk headers rag · how to improve chunk quality

How do I chunk PDFs, tables and code?

PDFs: extract with layout awareness so columns, headers and footers are handled, then split on headings and paragraphs, keeping page numbers on every chunk; PyMuPDF and Unstructured both expose structure [9][10]. Tables: keep each table whole, or split by row groups while repeating the header row in every chunk, and consider a text summary of the table as an extra chunk. Code: split on functions, classes and files, never mid-function, and include the file path and imports as context. Slide decks: one slide per chunk with the title.

Also asked as: how to chunk pdf for rag · chunking pdf documents · chunking tables rag · chunking code for rag · chunk markdown for rag · chunk html · chunking spreadsheets · chunking slides · pdf parsing for rag

In a packed college auditorium
In a packed college auditorium

What is parent-child or hierarchical chunking?

Index small chunks for precise matching, but store a pointer from each to a larger parent, the section or the page, and return the parent to the model at generation time. Retrieval is precise because small chunks have sharp embeddings; generation has context because the model reads the whole section. RAPTOR takes the idea further by building a tree of summaries over the corpus and retrieving at multiple levels [5]. For most systems, two levels, child and section, capture most of the benefit.

Also asked as: parent child chunking · hierarchical chunking rag · small to big retrieval · sentence window retrieval · raptor rag · multi level chunking · chunk hierarchy

How do I evaluate chunking?

Directly, at the retrieval stage, before any generation: build a set of real questions each labelled with the passage that answers it, and measure recall at k for your chunking configuration. Then read the misses. Was the answer split across two chunks? Was it retrieved but ranked low because the chunk was diluted? Was the heading lost? Each miss points at a specific chunking change. Generation metrics come later; you cannot generate a correct answer from a chunk that was never retrieved.

Also asked as: how to evaluate chunking · chunking evaluation metrics · retrieval recall chunking · rag evaluation chunking · test chunking strategy · measure chunk quality

What are common chunking mistakes?

Copying a chunk size from a tutorial. Splitting mid-sentence, mid-table or mid-function. Dropping the heading. Chunking a PDF's text in extraction order on a two-column page. Overlap so large that the reranker sees duplicates. No page or source metadata, so citations are impossible. Re-chunking without re-embedding. Never reading the retrieved chunks by eye. Every one of these produces a system that looks like it hallucinates and is actually retrieving garbage.

Also asked as: chunking mistakes · common rag mistakes chunking · why my rag retrieves wrong chunks · bad chunking examples · chunking pitfalls

Chunking interview questions

Explain why chunking matters more than the model. Compare fixed, recursive and structure-aware splitting. Say how you would choose a chunk size and defend the method, not the number. Explain parent-child retrieval. Describe how you would chunk a PDF with tables. Explain what metadata a chunk needs and why. These are stories if you have built a RAG system, definitions if you have not.

Also asked as: chunking interview questions · rag interview questions chunking · text splitting interview · llm interview chunking

Where should I start?

Take one real document set, write twenty questions and mark their answers, and index the corpus three ways: fixed 512 tokens, structure-aware, and structure-aware with headings prepended. Measure recall at 10 for each. The third will win, and you will understand why in a way no article can teach. My longer guide and the experiment write-up are on the portfolio [11][12]. For a hands-on retrieval session at your college, email pranjulrathour41@gmail.com or use pranjulrathour.scult.in/invite.

Sources

  1. Lewis et al., Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks (2020)arxiv.org
  2. Gao et al., Retrieval-Augmented Generation for Large Language Models: A Survey (2023)arxiv.org
  3. Barnett et al., Seven Failure Points When Engineering a RAG System (2024)arxiv.org
  4. Liu et al., Lost in the Middle: How Language Models Use Long Contexts (2023)arxiv.org
  5. Sarthi et al., RAPTOR: Recursive Abstractive Processing for Tree-Organized Retrieval (2024)arxiv.org
  6. Anthropic, Introducing Contextual Retrieval (2024)anthropic.com
  7. LangChain text splitters documentationpython.langchain.com
  8. LlamaIndex node parsers and text splittersdocs.llamaindex.ai
  9. Unstructured: open-source document partitioninggithub.com
  10. PyMuPDF documentationpymupdf.readthedocs.io
  11. Chunking strategies for RAG, Pranjul Rathourpranjulrathour.scult.in
  12. RAG chunk size experiments, Pranjul Rathourpranjulrathour.scult.in
  13. RAG.NextUpgrad source code, Pranjul Rathourgithub.com
Pranjul Rathour
Pranjul Rathour
GenAI Engineer · Kanpur, Uttar Pradesh, India

GenAI engineer and AI product builder with 2+ years shipping production-grade AI systems: RAG pipelines, fine-tuned LLMs, hybrid retrieval and multi-modal apps across vision, speech and OCR, architected end to end from ingestion to deployment. Leads engineering for SCULT INDIA's 14-member team, founded the 500+ member TechVerse Enclave community and has mentored 200+ students. Three hackathon first prizes: Changethon 2025 (IIT Roorkee), Product Genesis at Vividhotsava 2025 (CSJMU Kanpur) and BYTEBATTLE (MeetKats).

Open to GenAI roles, hackathon judging, mentorship sessions and guest talks at colleges. Email pranjulrathour41@gmail.com.

Keep reading

More on RAG & retrieval

All RAG & retrieval guides →
Campus talks · hackathon judging · mentoring

Want this as a live session at your college?

Open to GenAI roles, hackathon judging, mentorship sessions and guest talks at colleges.

Pranjul Rathour, GenAI engineer in Kanpur