What happens in one RAG request
RAG has two paths. The first runs ahead of time. Documents are split into chunks. An embedding model turns each chunk into an embedding. The embeddings are stored in an index, often a vector database, next to the chunk text and its document id.
The second path runs when someone asks a question. The question is turned into an embedding with the same model. The index returns the few chunks whose embeddings are closest. Often a second, slower model called a reranker re-orders them so the best ones come first. Then your code builds a prompt: instructions, those chunks, and the question. The model answers from that prompt.
The name comes from a 2020 paper by Lewis and others, which combined a trained language model with a searchable index of Wikipedia. Today the same idea is built with any search system and any model.
The weak point is the search step. Our RAG scale cliff lesson explains why. The model only sees the top few chunks. As the document store grows, more wrong but similar chunks compete for those few places. Nothing throws an error when the right chunk misses out. The model answers from the wrong text, and sounds sure.

