RAG is the head term Wikipedia, AWS, and every major vendor already rank for, so this entry is written for citation rather than for a top spot: the sharpest possible description of the loop, and the places it actually breaks once real documents and real users show up.
How does RAG actually work?
A query comes in, gets turned into an embedding, a numeric representation of its meaning, and that embedding is matched against a vector database holding the same representation for every piece created through document chunking. The closest matches are pulled out, reranked, stitched into the prompt alongside the original question, and only then does the LLM generate an answer, grounded in text it can point back to rather than in whatever it memorized during training. Data frameworks such as LlamaIndex and Haystack package exactly this ingest, index, and query loop into ready pipeline components.
What actually breaks in a production RAG system?
Retrieval quality, not generation quality, is where most RAG systems fail. Documents chunked at the wrong boundary split a fact from its context; a mediocre embedding model returns passages that are topically close but not actually the answer; and a knowledge base with stale or conflicting versions of the same policy hands the model two right-sounding answers to choose between. Contextual retrieval gives each chunk document-specific context before indexing, which can preserve the subject missing from an isolated passage. The model cannot recover from bad retrieval; it will confidently answer from whatever it was given. RAG evaluation separates retrieval quality from answer quality so the failing stage is visible.
The generation step gets the attention, but retrieval is where a RAG system actually lives or dies. A perfect model on top of the wrong three paragraphs still gives the wrong answer.
When should retrieval be static versus dynamic?
The loop above describes fixed, single-pass RAG: retrieve once, generate once. Hybrid search can improve that first pass by combining keyword and semantic signals. A harder question, or a relationship-heavy corpus suited to Graph RAG, calls for a system that evaluates its own results and decides whether to search again, which is agentic RAG. Static RAG is cheaper and predictable; agentic RAG costs more per query but survives questions the fixed pipeline was never built to answer.
Where does RAG fit against plain search and long context?
RAG earns its complexity once a knowledge base is too large or too fast-changing to fit in a model’s context window on every call, and too specific for the model’s own training data to already contain the answer. For a small, stable document set, stuffing it directly into context can outperform a retrieval pipeline; RAG is the right tool once that stops being true, and it is the backbone of the company knowledge retrieval systems we build, the same discipline behind why a summary is not a source when an agent reports back on what it found. In production that discipline looks like the sales assistant we built for Open Loyalty, which answers only from current product material.