What Is Retrieval-Augmented Generation (RAG) and When Should You Use It?

Ask a general-purpose language model what your company's returns policy says and you may well get a confident answer that is wrong in small, expensive ways. The model has never seen your documents. It knows only what it absorbed during training, plus whatever you paste into the prompt. Retrieval-augmented generation, usually shortened to RAG, fixes that by turning the model into a reader rather than an oracle. Before it answers, it goes and fetches the relevant passages from your own content — policy documents, product manuals, support tickets, contracts — and writes its reply from those. It is a modest idea with an outsized practical payoff.
Retrieval first, generation second
RAG is two systems stitched together. The first is a search layer over your material: anything from a vector database to Elasticsearch, Postgres full-text search or a well-tuned keyword index. The second is the language model that writes the answer. Neither is new. The value comes from the order in which they run.
- The user asks a question in ordinary language.
- The question is converted into a search query, often as a numeric embedding, sometimes alongside keywords.
- The retriever pulls back the handful of passages most likely to contain the answer.
- Those passages are placed in the model's prompt, together with the question and a strict instruction: answer using only this material, and say when it is not covered.
- The model writes a reply, ideally citing which document each part came from.
Where the passages come from
Your documents are split into chunks — a section, a page, a few paragraphs — and each chunk is indexed. Good chunking respects structure: headings stay with the text beneath them, tables are not sliced mid-row, and a warranty clause keeps its product name. Metadata matters just as much. If every chunk carries a product, version, region or department tag, you can filter before you search, which is usually cheaper and more accurate than relying on similarity alone. Most production systems use hybrid retrieval: keyword matching for part numbers and acronyms, semantic matching for the fuzzy way people actually phrase things.
Why this often beats fine-tuning
Fine-tuning adjusts how a model behaves — its tone, its formatting, its willingness to follow a house style. It is a poor way to store facts. Facts change. A price list, a holiday policy or an API endpoint is out of date the moment someone edits the intranet, and a fine-tuned model will keep reciting the old version until you retrain it. With RAG, updating the knowledge base means updating a document.
There is a second advantage that matters in any organisation with more than a handful of staff: permissions. Because retrieval happens at query time, you can filter results according to who is asking. A junior account manager searching the same index as a finance director gets a different set of passages, and therefore a different answer. Replicating that with a single fine-tuned model is close to impossible.
When RAG is the right tool
It fits best when the job is answering questions from a body of text that you own. Useful signals:
- Traceability is required. Someone needs to click through to the source and check it.
- The corpus is too big for a prompt and too expensive to paste in on every request.
- Content changes regularly. Policies, catalogues, technical bulletins, case notes.
- Access control matters. Different users should see different documents.
- The documents already exist in reasonable shape, as text rather than blurry scans.
- People ask in free text about material that is scattered across folders and systems.
When it is not
RAG is not a general fix for every language task. If you need an email rewritten, a support ticket classified or a long report summarised, a well-written prompt will usually do the job without any retrieval at all. If your whole knowledge base is a single stable page, put that page in the prompt and move on.
Be cautious with numbers that must be exactly right — stock levels, account balances, live delivery times. Those belong in a database query, with the model calling that query as a tool rather than guessing from prose. And if your documents are contradictory, outdated or poorly written, RAG will surface the mess faster and more convincingly. Fixing the source material comes first.
Where money, health or legal matters are involved, treat any generated answer as a first draft for a qualified person to review. RAG reduces invention; it does not remove the need for professional advice.
A worked example
Picture a mid-sized heating installer with a support inbox and around four hundred pages of manuals, warranty terms and service bulletins. Engineers on site ask things like "what's the flue clearance on the 30kW combi?" The team indexes the documents, tags each chunk with product model and document date, and builds a small interface with a search box. The retriever finds the relevant table, the model answers in two sentences and links to the page.
The same pattern works for an HR team fielding questions about parental leave, a charity explaining grant eligibility, or a software company whose API documentation is accurate but hard to search. The shape rarely changes.
Practical tips before you build
- Write down thirty real questions and their correct answers before writing any code. This is your test set.
- Check retrieval before you blame the model. Most bad answers are passages that were never found.
- Chunk by meaning, not by character count.
- Add metadata and use it to filter, especially for versioned or regional content.
- Instruct the model to quote and to say "not covered" rather than filling gaps.
- Show sources in the interface. People correct inaccuracies when they can see them.
- Log queries and retrieved passages, then retire documents nobody looks at.
- Cache common questions; it cuts both cost and waiting time.
A sensible way to start
Pick one audience, one document set and one high-volume question type. Build the thinnest version that answers those well, measure it against your test set, and only then widen the scope. RAG rewards patience with the boring parts — clean text, sensible chunks, honest evaluation — far more than it rewards a clever model choice.
Photo: felix_w / Pixabay


