New & Emerging

Effective RAG Chunking Strategies

Last updated 2026-09-19

Key points

What it is

  • **RAG (retrieval-augmented generation)** is an AI technique where an AI agent (a software that performs tasks) retrieves information it doesn't know from stored data.
  • **Chunking** is splitting documents into smaller pieces (called chunks) to prepare information for the AI to search and retrieve.
  • Chunking is like lossy compression (a way to reduce data size by discarding some information), meaning you might lose some details when you split documents.

How to use it

  • Chunk text by section, generate embeddings (numeric vectors capturing meaning) for each chunk, create a vector store (a database of those vectors), and add each embedding.
  • Generate an embedding for the user's question and search the vector store to find the most relevant chunks.
  • Expect real gains from revisiting and optimizing your chunking strategy, as it can close a 20-40% performance gap.

Watch out for

  • Chunks might not contain all the context needed to answer a question, and there's a hard limit on how much text you can feed the AI model.
  • If chunks are too big, you lose nuance and meaningful embeddings; if too small, you lose the big picture and efficiency.
  • Don't treat chunking as a one-time setup or a "fire-and-forget" task, as it's crucial for the AI's performance.

Tools named

  • Gemini File Search (a tool that simplifies RAG and chunking), Claude API (a tool for building AI applications)

Lesson 1: What is Effective RAG Chunking Strategies and why it matters

Effective RAG chunking strategies are the decisions you make about how to break documents into smaller pieces before storing them for an AI system to search later. RAG (retrieval augmented generation) is the concept where an AI agent only knows so much from its training data, so if you ask it something it doesn't know, it has to go find that information. Chunking is how you prepare that information. Yuval Belfer from AI21 Labs describes the typical approach: on day one, or week one, you pick a chunk size like 512, probably add some overlap of 10 or 20 percent, index everything, and forget all about it. That fixed strategy has real problems. If you chunk something too big, you get the whole thing back even when only part matters. Some people now say chunking is dead because agentic search tools have arrived, but those tools still don't solve everything. The payoff of doing chunking well is a repeatable mental model you can apply in your next implementation step. It matters because every later stage of your pipeline depends on it. You need a strong model for intensive tasks like extracting information from many chunks and analyzing on top of them. If you want to skip the infrastructure work entirely, Gemini File Search abstracts away the boring parts of RAG and lets you focus on the application rather than specific chunking strategies. But if you need control, you have to make deliberate choices.

Sources

Lesson 2: How to use Effective RAG Chunking Strategies: step-by-step

Chunking (splitting documents into pieces) is the step beginners skip, and it matters more than it looks. In a typical RAG (retrieval-augmented generation) system, you have two stages. The first is the boring one: picking a chunk size, often 512, adding 10-20% overlap, indexing everything, and forgetting it. Yuval Belfer argues this "Stop Chunking Like It's 2022" approach no longer holds up and that chunking isn't dead, despite claims that agentic search (AI agents that search on their own) replaces it. Agentic tools like grabs, LS, and finds help, but they aren't enough when you have lots of data and varied queries. The fix is a concrete pipeline. As the Claude API walkthrough shows, you chunk source text by section, generate embeddings (numeric vectors capturing meaning) for each chunk, create a vector store (a database of those vectors), and add each embedding. Then you generate an embedding for the user's question and search the store to find the most relevant chunks. The important move is to name the input, show where the application acts, and make the resulting state inspectable, so you keep a repeatable mental model. Step by step: chunk text by section, generate embeddings, create a vector store, embed the query, search. One caution: chunks might not contain all the context, and there's a hard limit on how much text you can feed the model. Finally, remember the answer to "is RAG dead?" depends entirely on your data.

Sources

Lesson 3: Best practices and pitfalls

Yuval Belfer argues in "Stop Chunking Like It's 2022" that most people treat chunking (splitting documents into smaller pieces) as a one-time setup: you pick a size like 512 tokens, add 10-20% overlap (repeated text between chunks), index everything, and forget it. He claims there is no single right chunk size, because chunking is essentially lossy compression (discarding information you can't recover). If chunks are too big, you lose nuance and meaningful embeddings (numeric representations of meaning); if too small, you lose the big picture and efficiency.

He notes that many now declare RAG dead and say chunking is dead too, because people use agentic search (AI that retrieves step by step) tools like grabs, LS, and finds. But he insists those aren't enough for large data and varied queries. His research found no chunk size dominates across datasets, yet better chunking strategy alone can close a 20-40% performance gap. The problem is that chunking is the "boring" stage few want to optimize. He also admits their tested sizes — 50, 100, 200 — were arbitrary, so future work must determine how many chunk sizes to use and which ones.

The Claude API lesson reinforces the pipeline: you chunk source text, generate embeddings, store them in a vector database (a store for similarity search), then match a user query against them. The practical lesson: stop treating chunking as fire-and-forget, and expect real gains from revisiting it.

Sources