Most retrieval-augmented generation systems that "hallucinate" are not hallucinating. They are faithfully summarising the wrong three paragraphs. The model did its job; retrieval handed it the wrong evidence.
This distinction matters because the two failures have completely different fixes, and teams routinely spend weeks swapping embedding models when the actual problem is that their chunker split a table away from its header.
The first thing worth establishing is which half of the pipeline is broken. Run your failing queries and look at what retrieval actually returned, before generation touched it.
Two outcomes, two very different projects:
- The right passage was retrieved and the answer is still wrong. That is a generation problem. Look at your prompt, your context ordering, and whether you are burying the evidence in the middle of a long context.
- The right passage was never retrieved. That is a retrieval problem, and no amount of prompt engineering will save you. The model cannot cite what it never saw.
In our experience the second case is far more common, and it is the one teams misdiagnose.
The default in almost every tutorial is to split on a fixed token count with some overlap. It is easy to implement and it is wrong for most real documents.
Consider a services agreement. Fixed-size chunking at 512 tokens will cheerfully produce a chunk that begins mid-sentence in clause 7.2 and ends partway through 7.4. Neither the clause number nor the defining context survives. When someone asks "what is the termination notice period", the chunk containing "thirty (30) days" has no indication that it is about termination.
The embedding is a faithful representation of a fragment that means nothing on its own.
Chunk on the document's own boundaries
Documents already tell you where the seams are. Use them:
- Markdown and HTML: split on headings, and carry the heading path down into each chunk so a section knows what it belongs to.
- Contracts and policies: split on numbered clauses.
- Transcripts: split on speaker turns or topic shifts, not on time.
- Tables: keep the header row with every chunk of the body, or serialise each row as a sentence.
Only after structural splitting should you enforce a size ceiling, and when a section genuinely exceeds it, split on paragraph boundaries rather than tokens.
| Strategy | Retrieval recall @5 | Answers citing the wrong clause |
|---|---|---|
| Fixed 512 tokens, 50 overlap | 61% | 22% |
| Clause-aware, heading path prepended | 94% | 3% |
The embedding model was never the bottleneck.
A chunk gets embedded in isolation, so it has to make sense in isolation. Two cheap techniques do most of the work:
Prepend the breadcrumb. If a chunk lives under "Master Services Agreement → 7. Termination → 7.2 Notice", put that path at the top of the chunk text before embedding. It costs a handful of tokens and dramatically improves matching on queries that name the topic.
Resolve the pronouns. A chunk that opens with "It shall not exceed thirty days" is unmatchable. A cheap rewriting pass that turns it into "The notice period shall not exceed thirty days" makes it findable. This is worth doing at index time, once, rather than at query time on every request.
None of the above is worth doing if you cannot tell whether it helped. You need a fixed evaluation set, and it does not need to be large.
Thirty to fifty real questions is enough to start. For each one, record the question, the passage that should be retrieved, and the correct answer. Pull them from actual support tickets or user logs rather than inventing them, since invented questions are always phrased more helpfully than real ones.
Then measure the two halves separately:
- Retrieval recall @k: how often the correct passage appears in the top k. This is the ceiling on everything downstream.
- Answer correctness, evaluated only on queries where retrieval succeeded. This isolates the generation half.
When recall @5 is 61%, no prompt change can push end-to-end accuracy past 61%. Knowing that saves you from optimising the wrong thing for a month.
Retrieval returning confident nonsense?
We run a two-week diagnostic on production RAG systems: golden set, recall baseline, and a ranked list of fixes with expected impact. Most of the wins are in chunking.
See how we workPull twenty queries your system got wrong last week. For each, check whether the correct passage was in the retrieved set at all. If it usually was not, stop tuning prompts and go look at your chunker. If it usually was, the problem is downstream and chunking is a distraction.
Either way you will know within an afternoon which half of the system to fix, which is considerably better than the alternative of changing both and hoping.
