Listen to this lesson
Unlock audio and more
Audio streaming, downloadable PDFs and certificates come with Plus and Pro.
What it means
A language model on its own answers from what it absorbed during training — which is frozen at some past date, contains nothing about your organization, and cannot be cited. Retrieval-augmented generation fixes all three problems with the same move: before answering, search a document collection for passages relevant to the question, paste those into the prompt, and ask the model to answer from them.
The pipeline is search, then generate. Documents are split into chunks, converted into embeddings, and stored in a vector database. A question is embedded the same way, the closest chunks are retrieved, and the model answers using that material — usually with citations back to the sources.
Concretely: an employee asks "how many vacation days carry over?" The system embeds that question, finds the three passages in the staff handbook whose meaning sits closest to it — which may never use the word "vacation" — and sends the model those passages plus the question. The model reads them and answers from what it was handed. Nothing about the handbook was learned; it was supplied, used once, and forgotten.
Everything upstream of the model decides quality. If the retrieval step returns the wrong three passages, a better model will simply write a more fluent wrong answer.
Why it matters
RAG is how most real business AI gets built. It is what lets a model answer questions about your contracts, your policies, or last week's numbers without retraining anything, and it is dramatically cheaper and faster to update than fine-tuning — new documents are available the moment they are indexed.
It also makes answers checkable. A grounded answer comes with sources a human can verify, which converts an unauditable assertion into something you can actually govern.
The permissions property is underrated and often decisive. Because retrieval happens per question, it can respect the asker's access rights — the same assistant answers differently for someone in HR and someone in sales, because the retrieval step never hands them documents they aren't cleared to see. A model that had absorbed the same documents during training could not make that distinction; the knowledge would be baked in for everyone.
What people get wrong
That RAG teaches the model your data. It does not. Nothing is learned and no weights change — the documents are pasted into the prompt at question time and forgotten the moment the answer is returned. Ask the same question with retrieval switched off and the model knows nothing about your business. This is why RAG updates instantly when a document changes, and why it is not a substitute for fine-tuning when you need different behavior rather than different facts.
That a big context window makes RAG obsolete. Windows have grown enough to hold entire document sets, so the argument goes that you can skip retrieval and paste everything in. Three things push back: you pay per token on every request, so sending a whole corpus each time is expensive; latency scales with what you send; and models attend unevenly across a long context, so material buried in the middle gets less consideration than the same material retrieved and placed deliberately. Retrieval is a relevance filter, and relevance filters do not stop being useful because the pipe got wider.
That RAG eliminates hallucination. It reduces it and — more importantly — makes it checkable. A model can still misread a retrieved passage or cite a source that does not actually support the claim. What changes is that a reader can follow the citation and see.
In practice
When RAG disappoints, the failure is nearly always retrieval rather than generation: the right passage was never fetched, so the model never had a chance. Debug by looking at what was retrieved before blaming the model or reaching for a bigger one.
Test the abstention path explicitly — ask something your sources genuinely do not cover and see whether the system says so or fills the gap from training. That single test tells you more about whether it is safe to deploy than any benchmark.
Where this shows up
Tools and models in our catalog.
PineconeThe leading managed vector database for AI applications. Serverless pricing, 99.99% SLA, and billions of vectors at millisecond query speeds. Widely used in production RAG systems.
Supabase VectorPostgreSQL-based vector storage using the pgvector extension. Seamlessly combines traditional relational data with vector search in a single database.
LlamaIndexOpen-source framework specialized for building RAG (Retrieval-Augmented Generation) systems and data-aware LLM applications. Strong for enterprise knowledge bases.
LangChainThe dominant open-source framework for building LLM-powered applications and agents. Python/JS libraries plus LangSmith for tracing and LangGraph for complex agents.