Skip to content

Courses

Learning streak: 6 days in a row

Reading · 9 min · Lesson 3 of 7

Retrieval quality

Most bad RAG answers are created before the model writes a word.

When a RAG assistant gives a wrong answer, the instinct is to blame the model. Usually the problem is retrieval: the right passage was never found, or it was outranked by an outdated version of the same document. A model given the wrong context will confidently answer the wrong question.

Measure retrieval separately from answers. Build a test set of 30 to 50 real questions, mark the passage that should answer each one, then check how often that passage appears among the top results. This single number tells you more than any demo.

The cheapest improvements are usually in the content: remove duplicates and old versions, add clear titles and dates, and split long mixed documents into focused ones. Clean sources do more for accuracy than switching to a bigger model.

Key takeaways

  1. Wrong answers usually start with wrong retrieval.
  2. Test retrieval with real questions and known answers.
  3. Cleaning the content often beats upgrading the model.