Retrieval-augmented generation (RAG) is an architecture that retrieves relevant passages from a knowledge source—documents, FAQs, product data, policies—and provides them to a generative model before it answers. The model still writes the response, but it is steered by retrieved evidence rather than relying only on training-time memory.
RAG is one of the most practical ways to make generative AI useful for company-specific questions. Without retrieval, fluent answers can still be confidently wrong about your prices, SOPs, or warranties.
Why RAG exists
Models alone do not automatically know your latest price list, SOP, or warranty policy. Fine-tuning is not always the right tool for frequently changing facts. RAG keeps knowledge in searchable stores you can update, then fetches snippets per question.
RAG answer path
- User asks a question
- Retriever finds relevant chunks
- Prompt includes question + evidence
- Model drafts an answer from evidence
- Optional citations are shown to the user
What makes RAG quality rise or fall
- Clean source documents with clear ownership
- Sensible chunking and metadata (product, language, date)
- Permission-aware retrieval so users only see allowed content
- Evaluation sets that catch hallucinations and misses
- Fresh indexing when documents change
- An explicit “I don’t know / escalate” behaviour when evidence is weak
RAG vs agents
RAG grounds answers. Agents take actions. Many production systems combine both: retrieve policy, then call a tool if the user is allowed to proceed. Do not confuse a grounded chatbot with a tool-using agent—they solve different jobs.
Chunking and metadata strategy
Chunk by meaningful sections (headings, product SKUs, policy clauses). Attach metadata such as brand, language, audience, and effective date so retrieval can filter before ranking. Bad chunking produces answers that mix obsolete and current policy.
Citations users can trust
Show which document and section supported an answer. If the model cannot find evidence, it should say so and escalate—not invent a confident policy. Citations also help content owners fix the source instead of endlessly tweaking prompts.
Permissions
Retrieval must respect ACLs. A staff handbook answer must not surface executive-only compensation files. Permission-aware indexes are part of RAG, not an add-on after a scary incident.
A concrete example
An employee asks, “What is our return window for refurbished laptops?” RAG retrieves the current returns policy section for refurbished goods, the model answers with that window, and the UI cites the policy document. When the policy updates next quarter, you re-index the document—you do not retrain a model to “remember” the new number.
Operating the knowledge base
Assign owners per content domain. Stale PDFs in a shared drive become stale answers. Treat indexing freshness, document retirement, and review cadences as product operations—the same seriousness you give CRM data quality.
FAQ
RAG or fine-tuning?
Prefer RAG for changing factual knowledge. Consider fine-tuning for style or specialised formats after you have evaluation discipline—not as the first hammer.
Will RAG eliminate hallucinations?
It reduces them when evidence is strong and prompts require grounding. It does not make models incapable of inventing. Evaluation and escalation remain mandatory.
What content should we index first?
High-traffic FAQs, current policies, product specs, and approved price rules. Skip obsolete folders and personal drafts until governance exists.
Can customers use the same RAG as staff?
Only with separate permission scopes and content sets. Customer-facing indexes should exclude internal playbooks and confidential pricing exceptions.
How does RAG help chatbots and automation?
It turns AI chatbots and business automation into systems that answer from your knowledge—not from generic internet memory.
Related concepts
Models: what is generative AI. Actions: what is an AI agent. Channel choice: AI chatbot vs traditional chatbot. Process fit: AI for business automation.