Video · 9 min · Lesson 1 of 7
RAG in one loop
Question, retrieval, answer: how a model gets your knowledge without retraining.
Transcript
A general model knows the public internet up to its training cut-off, but not your pricing rules, your procedures or last week's decision. Retrieval-augmented generation closes that gap with a simple loop: a user asks a question, the system retrieves the most relevant passages from your documents, and the model answers using those passages as its source.
Nothing is retrained. Your documents stay in your own index and are handed to the model only as context for each question. That makes RAG faster to build than fine-tuning, easy to update when a document changes and much easier to audit, because every answer can point to the text it came from.
Fine-tuning still has a place when a model needs to learn a style or a narrow format. But when the goal is answering from facts that change, RAG is almost always the better first choice.
Key takeaways
- RAG retrieves relevant passages and answers from them.
- Your knowledge stays in your index; the model is not retrained.
- For facts that change, RAG usually beats fine-tuning.