A language model can write an answer from what it learned during training. Retrieval-augmented generation, usually shortened to RAG, adds another step: find relevant information from your own documents or database, then give that context to the model before it responds.
Business knowledge is different from model knowledge
A support assistant may need your refund policy. An internal assistant may need the latest project handbook. A product chatbot may need documentation that changes every week. These are not things you want the model to guess from general knowledge.
RAG is a good fit when the answer lives in a body of information that your team owns and updates. It lets the system search that information at the moment a question arrives.
Good situations for RAG
- Customer support over help centre and policy content.
- Internal questions over handbooks, procedures and project documents.
- Technical documentation where answers should point back to source material.
- Product or catalogue questions where the underlying data changes regularly.
When RAG is unnecessary
If the assistant only needs to classify a request, follow a fixed flow or answer from a handful of stable rules, conventional software may be better. A search box, decision tree or small set of curated responses can be faster, easier to test and easier to explain.
RAG also does not make a chatbot knowledgeable by magic. Poor documents, vague chunks and weak retrieval produce poor context. The model can still misunderstand that context or confidently answer beyond it.
Treat retrieval as a product problem
Good RAG work includes deciding what content is trustworthy, how it is split, what metadata matters, how results are ranked and what happens when no useful result is found. The experience should make uncertainty visible rather than hiding it behind a polished sentence.
Permissions matter
A search layer must respect the permissions people already have. A helpful assistant that exposes a private document is not helpful. Filter retrieval by user, team or role before context reaches the model, and keep an audit trail for sensitive workflows.
Useful work usually starts with a clearer question.