Tailr
← All posts

What Is RAG? Retrieval-Augmented Generation Explained Simply

· updated

How RAG works: a question is used to search your documents, the most relevant passages are added to the prompt, and the model answers with citations

Ask a chatbot about your company’s leave policy and it will either admit it doesn’t know or, worse, invent something plausible. The model was never trained on your documents. RAG is the standard fix, and it’s behind most “chat with your docs” features, AI support bots and internal company assistants you’ve used. This guide explains what RAG is, how it works step by step, when to use it instead of fine-tuning, where it goes wrong, and how to build a simple version yourself.

The short answer: RAG (retrieval-augmented generation) is a technique that makes a language model answer using information retrieved from your own documents, instead of relying only on what it learned in training. It works in four steps:

  1. Index: split your documents into chunks and store them in a searchable index, usually as embeddings in a vector database.
  2. Retrieve: when a question comes in, search the index for the most relevant chunks.
  3. Augment: add those chunks to the prompt, along with instructions such as “answer only from this context and cite your sources”.
  4. Generate: the model writes an answer grounded in the retrieved text.

RAG keeps answers current (update the documents, not the model), reduces made-up answers and allows citations. The term comes from a 2020 paper by Patrick Lewis and colleagues at Facebook AI Research.

Why RAG exists: the problem it solves

A large language model knows what was in its training data, up to a cutoff date. It doesn’t know:

  • Your private information: company policies, product manuals, support tickets, contracts, your own notes.
  • Recent information: anything after its training cutoff.
  • Exact details: it may remember the gist of a document but not the precise clause, number or date.

When asked about these, a model can hallucinate: produce a fluent, confident, wrong answer. You could retrain the model on your documents, but that’s slow, expensive, and has to be repeated every time the documents change. RAG takes a simpler route: find the relevant text at question time and hand it to the model.

Think of it as an open-book exam. Instead of hoping the student memorised the textbook, you let them look up the right page before answering.

How RAG works, step by step

Step 1: prepare and chunk the documents

Collect the source material: PDFs, web pages, docs, tickets, database rows. Clean it (remove navigation menus, headers, junk) and split it into chunks, typically a few hundred words each, often with a little overlap so ideas aren’t cut in half. Chunk size matters: too small and chunks lose context; too large and they carry noise.

Step 2: create embeddings and store them

Each chunk is turned into an embedding: a list of numbers that represents its meaning, produced by an embedding model. Chunks about “annual leave” and “vacation days” end up with similar embeddings even though the words differ. The embeddings, along with the original text and metadata (source, page, date), go into a vector database or search index.

Step 3: retrieve relevant chunks

When a user asks a question, it’s embedded the same way, and the system finds the chunks whose embeddings are closest in meaning. Many production systems use hybrid search, combining meaning-based (vector) search with classic keyword search, because keyword search is better at exact terms like product codes and names. A reranker model often re-scores the top results to put the most useful ones first.

Step 4: augment the prompt

The top few chunks are inserted into the prompt with clear instructions, for example:

Answer the question using only the context below. Cite the source of each
fact as [source, page]. If the answer isn't in the context, say you don't
know.

Context:
[chunk 1 — HR Policy.pdf, p. 4]
[chunk 2 — HR Policy.pdf, p. 5]
[chunk 3 — FAQ page, updated 2026-08]

Question: How many days of annual leave do new employees get?

Step 5: generate the answer

The model writes an answer grounded in the retrieved context, with citations the user can click to check. Good systems also return “I don’t know” when retrieval comes back empty or irrelevant.

RAG vs fine-tuning vs long context

RAG Fine-tuning Long context (paste it all in)
What it changes What the model sees at question time The model’s weights What the model sees at question time
Best for Facts that change, private knowledge, citations Consistent style, format or specialised behaviour Small document sets, one-off analysis
Updating knowledge Update the documents Retrain Paste the new version
Cost Moderate Higher upfront Grows with every question
Citations Natural Hard Possible

They combine well: a fine-tuned model can still use RAG, and a RAG system can retrieve generously when the model has a long context window. What is fine-tuning covers when tuning is worth it.

Real-world examples of RAG

  • Customer support bots that answer from the help centre and cite the article.
  • Internal assistants that answer HR, IT and policy questions from company documents.
  • Legal and compliance research across contracts and regulations, with clause-level citations.
  • Developer documentation search that answers “how do I…” with the right code snippet.
  • Sales and customer success tools that pull answers from past proposals and product specs.
  • “Chat with your PDF” features in many apps.

Where RAG goes wrong

Most RAG failures are retrieval failures, not model failures:

  • Bad chunking: a table split across chunks, or an answer that needs two sections that were separated.
  • Missed retrieval: the question uses different words than the document, and there’s no keyword fallback.
  • Wrong or stale documents: old versions retrieved alongside new ones.
  • Too much context: ten loosely related chunks bury the one that matters.
  • The answer isn’t there: and the model answers anyway because the prompt didn’t allow “I don’t know”.
  • Prompt injection: a retrieved document contains instructions (“ignore previous rules…”) that the model follows.
  • No evaluation: nobody measures accuracy, so quality quietly degrades.

The fix for most of these is the same habit: build an evaluation set of real questions with known answers, and measure retrieval and answer quality every time you change something.

How to build a simple RAG app

A small RAG project is one of the best portfolio pieces for AI engineering, data and developer roles. The outline:

  1. Pick a document set with real questions: a university’s public policies, a product’s docs, a government scheme’s FAQs.
  2. Chunk and embed the documents with an embedding model, and store them in a vector store (many are free to run locally).
  3. Write the retrieval step: embed the question, fetch the top five chunks, optionally rerank.
  4. Write the prompt with the context, citation instructions and permission to say “I don’t know”.
  5. Call a model API, such as the Claude API, and show the answer with its sources.
  6. Build an eval set of 25 questions with expected answers, and a script that scores the system.
  7. Improve and measure: change chunk size, add hybrid search or a reranker, and show the before and after numbers.

You can build a version of this in a weekend with Claude’s help. Weekend projects to build with Claude to get hired includes a project brief and prompt for exactly this, and how to use Claude Code covers the tool.

RAG, agents and where it’s heading

RAG started as a fixed pipeline: retrieve once, then answer. Increasingly it’s one tool inside an AI agent: the agent decides when to search, rewrites the query if the first search fails, searches several sources, and checks the answer before replying. This is sometimes called agentic RAG. The fundamentals don’t change: good chunks, good retrieval, clear instructions and an evaluation set.

RAG terms, quickly

  • Chunk: a piece of a document stored for retrieval.
  • Embedding: numbers that represent a chunk’s meaning.
  • Vector database: stores embeddings and finds the closest ones.
  • Hybrid search: vector search plus keyword search.
  • Reranker: a model that re-orders retrieved chunks by relevance.
  • Grounding: making the answer rely on retrieved sources.
  • Context window: how much text the model can read at once. See context engineering vs prompt engineering for why what goes in it matters.

RAG and your career

RAG is one of the most common topics in AI engineering interviews and job listings. If you’re aiming for these roles, AI engineer interview questions covers what gets asked, and how to get started with AI engineering gives a learning path.

When you apply, the listings vary: some want RAG and enterprise search, others want agents or fine-tuning. Lead with the experience each one asks for. Tailr is a Chrome extension that tailors your resume to the job listing you’re viewing, using only your real experience, then writes a cover letter and tracks the application.

Try Tailr

Conclusion

RAG is an open-book exam for AI: before answering, the system looks up the most relevant passages from your documents and gives them to the model, which answers from them and cites its sources. It keeps answers current, reduces made-up facts, and doesn’t require retraining. Most of the quality comes from the unglamorous parts, chunking, retrieval and evaluation, which is exactly why building one is such a good way to learn how AI products really work.

Frequently asked questions

01What is RAG in simple terms?

RAG, or retrieval-augmented generation, is a way of making an AI model answer from your own documents. Before the model answers, the system searches a set of documents for the passages most relevant to the question and adds them to the prompt, so the answer is based on that information rather than only on what the model learned in training.

02Why is RAG used?

RAG lets a language model use information it was never trained on, such as company policies, product docs, or recent data, without retraining the model. It reduces made-up answers, keeps answers current because you just update the documents, and lets the model cite its sources so people can check them.

03What is the difference between RAG and fine-tuning?

RAG gives the model new information at question time by retrieving relevant documents; fine-tuning changes the model itself by training it further on examples. Use RAG when the model needs facts that change or are specific to you. Use fine-tuning when you need a consistent style, format or behaviour. Many teams start with RAG because it's cheaper and easier to update.

04What is a vector database in RAG?

A vector database stores embeddings, which are lists of numbers that capture the meaning of each chunk of text. When a question comes in, it's turned into an embedding too, and the database finds the chunks whose meaning is closest. This lets RAG find relevant passages even when they don't share exact words with the question.

05Does RAG stop AI hallucinations?

It reduces them but doesn't eliminate them. If retrieval finds the wrong passages, or the right answer isn't in the documents, the model can still produce a confident wrong answer. Good RAG systems tell the model to say it doesn't know when the context doesn't contain the answer, show citations, and are tested with an evaluation set.

06Is RAG still relevant with long context windows?

Yes. Models can now read hundreds of thousands of tokens, which makes it possible to paste in whole documents for small collections. But for large or frequently changing knowledge bases, retrieving only the relevant passages is still cheaper, faster, easier to keep up to date, and often more accurate than stuffing everything into the prompt.