Key takeaways
- RAG retrieves passages from your own documents first, then has the model answer from them, so the answer can cite sources and stay current without retraining.
- Use RAG for facts the model does not have; use fine-tuning for how the model behaves. Most products need RAG first.
- The pipeline is chunk, embed, index, retrieve, rerank, generate. Retrieval quality decides answer quality far more than the model does.
- Hybrid search (BM25 plus embeddings, fused with reciprocal rank fusion) and a cross-encoder reranker fix most bad answers.
- Measure retrieval recall and answer faithfulness separately, and add a confidence gate so the system says "I don't know" instead of inventing.
What is RAG in AI?
RAG, retrieval-augmented generation, is a way of building AI answers in two moves: first search a store of your own documents for the passages most relevant to the question, then hand those passages to a large language model and have it write the answer from them, with citations. The model no longer answers from memory alone. It answers from evidence you chose, which is why RAG is the default architecture for chatbots over company documents, support desks, legal and medical search, and any product where "I'm not sure" beats a confident guess.
Also asked as: what is rag in llm · what is rag in generative ai · what does rag stand for in ai · rag meaning in ai · what is retrieval augmented generation
The term comes from a 2020 paper by Patrick Lewis and colleagues at Facebook AI Research, which combined a dense retriever with a sequence-to-sequence generator and showed the pair beat purely parametric models on knowledge-intensive tasks while producing more specific, factual language [1]. The idea has since grown into an entire engineering discipline, catalogued in surveys such as Gao et al. [2].
The one-sentence version I give students: a plain LLM is a very well-read person answering with the library door locked. RAG unlocks the door, walks them to the right shelf, and asks them to quote the page.
A RAG demo works on the one document you tested it with. A RAG product works on the document a stranger uploads at 2 a.m. Everything in this guide is about closing that gap. Pranjul Rathour, from building RAG.NextUpgrad
How does RAG work?
RAG runs in two phases. Offline, you split documents into chunks, turn each chunk into an embedding vector and store it in an index. Online, you embed the user's question, pull the closest chunks, optionally rerank them, and prompt the model with the question plus those chunks so it answers from them.
Also asked as: how rag works in llm · how does rag work in ai · how retrieval augmented generation works · how does a rag model work · rag architecture explained
The offline phase is where most quality is won or lost, because everything downstream can only choose among the chunks you created. The online phase is where latency and cost live, since each question costs an embedding call, a search, possibly a reranker pass and a generation.
In RAG.NextUpgrad, the platform I built at NextUpgrad in 2026, the online path looks like this: the query is embedded and also tokenised for BM25, both retrievers return candidates, reciprocal rank fusion merges them, a cross-encoder reranks the top set, a confidence gate checks whether the evidence actually supports an answer, and only then does the model stream a response with inline citations [17].
What is a RAG pipeline?
A RAG pipeline is the chain of components that turns a raw document collection and a user question into a cited answer: loaders, a chunker, an embedding model, a vector index, usually a keyword index, a retriever, a reranker, a prompt template, the LLM and an evaluation harness. Frameworks such as LangChain and LlamaIndex give you each piece; the engineering is in choosing and tuning them for your documents.
Also asked as: what is rag pipeline · rag pipeline explained · best rag pipeline · how to build rag pipeline
What is a RAG model?
There is no separate "RAG model". RAG is a system pattern that wraps an ordinary language model, open or closed, with a retriever. The original paper trained the retriever and generator jointly [1], but almost every production system today uses an off-the-shelf embedding model for retrieval and an off-the-shelf LLM for generation, connected by code.
Also asked as: what is rag model · rag model in ai · is rag a model or a technique · rag llm meaning
Why do we need RAG? Why is it important?
RAG solves three problems a plain LLM cannot: it knows nothing after its training cutoff, it knows nothing about your private data, and when it lacks facts it invents plausible text. Retrieval supplies current, private, checkable facts at answer time, so the model can quote instead of guess.
Also asked as: why rag is important · why use rag in llm · why is rag needed · what problem does rag solve · benefits of rag
The practical benefits, in the order clients care about them:
- Grounding and citations. Every sentence can point at the passage it came from, which makes answers auditable. This matters more than raw accuracy in regulated work.
- Freshness without retraining. Add a document to the index and it is answerable a minute later. Retraining or fine-tuning a model for new facts takes hours and still does not cite.
- Access control. Retrieval can filter by user, tenant or clearance before the model ever sees the text. A fine-tuned model has no such boundary.
- Cost. Retrieval is cheap compared with training, and with a small model plus good retrieval you often match a much larger model answering from memory.
The limit is just as important: RAG helps with what the model knows, not how it reasons or writes. If your problem is tone, format or a specialised skill, retrieval will not fix it. That is where fine-tuning belongs.
What is the difference between RAG and fine-tuning?
RAG changes what the model can see at answer time; fine-tuning changes the model's weights, so it changes how the model behaves. Use RAG when the answer depends on facts, documents or anything that changes. Use fine-tuning when you need a consistent style, a strict output format, a niche skill, or lower latency and cost on a narrow task. Many production systems do both.
Also asked as: rag vs fine tuning · rag vs fine tuning vs prompt engineering · when to use rag vs fine tuning · difference between rag and fine tuning · is rag better than fine tuning · rag vs llm
I have built both kinds of platform in the same year, RAG.NextUpgrad and FineTune Studio, and the question I ask clients is one word: is the problem knowledge or behaviour? Knowledge is retrieval. Behaviour is training. A support bot that must know your refund policy is a knowledge problem. A model that must write incident reports in your house format is a behaviour problem. A bot that must do both is RAG for the policy and a small fine-tune for the voice.
What is the difference between RAG and prompt engineering?
Prompt engineering writes better instructions for the model; RAG adds a retrieval system that supplies evidence to those instructions. Prompt engineering alone cannot give the model facts it never saw, but every RAG system still needs a well-engineered prompt that tells the model to answer only from the passages and to say so when they are insufficient.
Also asked as: rag vs prompt engineering · rag vs context engineering · is rag prompt engineering
What is chunking in RAG, and what is the best chunking strategy?
Chunking is splitting documents into passages that can be embedded and retrieved on their own. The best strategy respects the document's structure, headings, paragraphs, table rows and code blocks, keeps chunks small enough to be precise and large enough to be understood alone, and prepends the section heading to each chunk so a passage carries its context.
Also asked as: what is chunking in rag · best chunking strategy for rag · best chunking size for rag · chunking techniques in rag · how to chunk documents for rag · rag chunk size
There is no universal chunk size, and anyone quoting one without asking about your documents is guessing. What holds across the projects I have shipped:
- Chunk for the question, not the document. A chunk should be able to answer a realistic user question by itself. For policy documents that is often a few hundred tokens; for API references it can be one function signature plus its description.
- Overlap is insurance, not a strategy. A small overlap catches sentences that straddle a boundary. Large overlaps duplicate what the reranker sees and inflate the index.
- Never split inside a structure. Tables, code blocks, numbered procedures, and a heading with its first paragraph belong together. Splitting them yields chunks that are individually meaningless.
- Carry the heading. Prepending "Section 4.2, Rate limits" to a chunk that says "the limit is 5 requests per second" did more for answer accuracy in RAG.NextUpgrad than any embedding model change.
Barnett et al. list missing content and wrongly ranked chunks among the seven failure points of RAG systems, and both usually trace back to chunking decisions [10].
What are embeddings, and which embedding model is best for RAG?
An embedding is a list of numbers, a vector, that represents the meaning of a piece of text so that texts about the same thing end up close together. RAG uses embeddings to find chunks that are semantically similar to the question even when they share no keywords. The best model is the one that scores well on retrieval for your language and domain on the MTEB benchmark and fits your latency and cost budget; test two or three on your own questions before committing.
Also asked as: what is text embedding · best embeddings for rag · best embedding model for rag · embeddings vs tokens · how do embeddings work in rag · what is vector embedding
Dense retrieval with learned embeddings was shown to beat traditional keyword retrieval for open-domain question answering by Karpukhin et al. in 2020 [3], and that result is why almost every RAG system starts with a vector index. Three practical notes:
- Dimensions are a trade-off, not a score. A 1,536-dimension vector costs more storage and search time than a 384-dimension one. Whether it retrieves better for you is an empirical question.
- Embeddings are not tokens. Tokens are the pieces a model reads; an embedding is a single vector summarising a whole chunk. People search for the difference often enough that it deserves saying plainly.
- Switching models means re-embedding everything. Vectors from two models are not comparable. Budget for a full re-index when you upgrade.
The MTEB leaderboard on Hugging Face ranks open and commercial embedding models on retrieval and other tasks [14]. Use it to shortlist, then evaluate on your own documents, because leaderboard corpora are not your corpus.
What is a vector database, and do I need one?
A vector database stores embeddings and answers "which stored vectors are closest to this one" quickly, usually with an approximate nearest-neighbour index. You need one when your index outgrows memory, needs filtering by metadata, or must serve many concurrent users. For a prototype or a few hundred thousand chunks, a library such as FAISS or the pgvector extension inside the Postgres you already run is enough.
Also asked as: what is vector database · what is vector database in ai · what is vector database in llm · best vector database for rag · do i need a vector database for rag · how vector databases work · faiss vs qdrant vs pgvector · what is vector search
FAISS is described in Johnson, Douze and Jégou's 2017 paper and remains the reference implementation for in-process search [7]. pgvector adds vector columns and indexes to PostgreSQL [15]. Qdrant is documented at qdrant.tech [16]. My default for a new client project is pgvector, because the vectors then live next to the metadata, the permissions and the audit log, and "one more database to operate" is a real cost for a small team.
What is hybrid search in RAG?
Hybrid search runs two retrievers on every question, a keyword retriever such as BM25 and a dense embedding retriever, then merges their ranked lists. Keywords catch exact terms, product codes, names and rare words that embeddings blur; embeddings catch paraphrases and meaning that keywords miss. Fusing them, most simply with reciprocal rank fusion, gives better recall than either alone.
Also asked as: what is hybrid search in rag · hybrid search vs vector search · bm25 vs embeddings · what is bm25 · reciprocal rank fusion rag · hybrid retrieval rag
BM25 is the classic probabilistic ranking function, described in full by Robertson and Zaragoza [4]. Reciprocal rank fusion, from Cormack, Clarke and Buettcher, scores each document by the sum of 1 divided by (a constant plus its rank) in every list, which needs no score normalisation and was shown to outperform more complex fusion methods [5]. It is a dozen lines of code and it is in RAG.NextUpgrad exactly because of that paper.
The pattern I see in support corpora: users ask with the product's own words ("error E-4021", "Form 16"), and dense retrieval alone often returns something semantically nearby but wrong. BM25 nails the exact token. The reverse happens for how-do-I questions phrased in plain language, where BM25 misses and embeddings win. Hybrid is not a luxury; it is the fix for both.
What is reranking in RAG?
Reranking is a second scoring pass over the top retrieved candidates using a model that reads the question and the passage together, a cross-encoder, and outputs a relevance score. Retrievers compare vectors computed separately for speed; a reranker looks at the pair, which is far more accurate, so you retrieve, say, 30 candidates and rerank to the best 5 before generation.
Also asked as: what is reranking in rag · why rerank in rag · cross encoder vs bi encoder · best reranker for rag · does rag need a reranker
Passage reranking with BERT-style cross-encoders was introduced by Nogueira and Cho in 2019 and remains the standard recipe [6]. It also addresses a known weakness of long prompts: Liu et al. showed that models use information at the start and end of a long context far better than information in the middle [8], so sending five well-ordered passages beats sending thirty in a heap.
How do I build a RAG chatbot?
Start with one document type and ten real questions. Load and chunk the documents with structure-aware splitting, embed with a model you shortlisted from MTEB, index in pgvector or FAISS, add BM25, retrieve with both and fuse with reciprocal rank fusion, rerank with a cross-encoder, and prompt the model to answer only from the passages and to cite them. Measure retrieval recall on your ten questions before touching the prompt. Then add streaming, a confidence gate and logging.
Also asked as: how to build rag chatbot · how to build a rag application · how to implement rag · how to create rag pipeline · rag tutorial for beginners · how to build rag with langchain · how to make a rag chatbot in python
The skeleton of the query side, without any framework, is short enough to read:
def answer(question: str) -> dict:
q_vec = embed(question) # same model used at index time
dense = vector_index.search(q_vec, k=30) # pgvector / FAISS / Qdrant
sparse = bm25_index.search(question, k=30) # keyword retriever
fused = reciprocal_rank_fusion([dense, sparse], k=60)
top = rerank(question, fused[:30])[:5] # cross-encoder scores the pairs
if confidence(question, top) < THRESHOLD: # evidence too weak
return {"answer": None, "reason": "no supporting passages", "sources": []}
prompt = build_prompt(question, top) # "answer only from these passages, cite [n]"
return {"answer": llm(prompt), "sources": [c.source for c in top]}
Everything else, ingestion jobs, a UI, authentication, tenant filters, is ordinary software engineering. In RAG.NextUpgrad that engineering includes Docker, CI/CD, structured logging, and a test suite of 196 tests, and the whole service runs in about 220 MB of RAM on a free tier, which is a useful reminder that RAG is not inherently heavy [17].

Which framework should I use, LangChain or LlamaIndex, or none?
Use a framework to prototype and to borrow loaders, then own the retrieval code yourself once you know what your system needs. Both LangChain and LlamaIndex are excellent for getting to a first answer in an afternoon. The trouble comes later, when a bug sits under three layers of abstraction. The query path above is a few hundred lines; you can read all of it.
Also asked as: langchain vs llamaindex for rag · best framework for rag · do i need langchain for rag · rag without langchain
How do I evaluate a RAG pipeline?
Evaluate retrieval and generation separately. For retrieval, build a set of real questions with the passages that should be returned and measure recall at k, the share of questions where a correct passage is in the top k. For generation, measure faithfulness, whether every claim in the answer is supported by the retrieved passages, and answer relevance. Frameworks such as RAGAS automate the generation metrics with an LLM judge, but the retrieval set you must build by hand.
Also asked as: how to evaluate rag pipeline · how to evaluate rag chatbot · rag evaluation metrics · what is faithfulness in rag · rag recall · ragas evaluation · how to test rag
RAGAS, from Es et al., defines faithfulness, answer relevance and context relevance as reference-free metrics scored by a model [9]. They are useful, and they are not a substitute for a human-labelled retrieval set, because a system can be perfectly faithful to the wrong passage.
The metric that changed my own systems most was the simplest: for each question, was the right chunk anywhere in the top 30? When it was not, no prompt or model could help, and the fix was always upstream, in chunking or in adding keyword retrieval.
Why does RAG fail, and how do I fix it?
RAG fails for a small number of repeatable reasons: the answer was never in the index, it was in the index but chunked so it lost its meaning, retrieval ranked the right chunk too low, the model ignored the passages and answered from memory, or the passages contradicted each other and the model picked wrong. Each has a specific fix, and the fix is almost never "a bigger model".
Also asked as: why rag fails · rag failure modes · rag hallucination · how to reduce hallucination in rag · rag problems · limitations of rag · challenges of rag
Barnett et al. catalogue seven failure points from three case studies, from missing content to incomplete answers, and their central observation matches my experience: you only discover most failures in operation, so logging every retrieval and every answer is part of the architecture, not an add-on [10].
What is a confidence gate, and why did I build one?
A confidence gate is a check between retrieval and generation that measures how strongly the retrieved passages support an answer, and refuses to generate when support is weak. Instead of a fluent hallucination the user sees "I could not find this in the documents", with the option to rephrase or upload more. It turns the most dangerous RAG failure, a confident wrong answer, into an honest non-answer.
Also asked as: what is a confidence gate in rag · how to make rag say i don't know · rag refusal · rag abstain · how to prevent rag hallucination
In RAG.NextUpgrad the gate looks at reranker scores and the agreement between passages before any tokens are generated. The threshold is tuned on the evaluation set so that it rarely blocks answerable questions. This is the single feature clients noticed most, and I wrote up the reasoning in a separate article [18]. Research is moving the same way: Self-RAG trains a model to decide when to retrieve and to critique its own output with reflection tokens [12].
The model is never allowed to guess. If the evidence is not there, the honest answer is that the evidence is not there, and the product says so out loud. Pranjul Rathour, RAG.NextUpgrad design note
What is agentic RAG, and what is Graph RAG?
Agentic RAG puts a planning loop around retrieval: the model decides whether to retrieve, rewrites the query, calls several retrievers or tools, checks the results, and retrieves again if needed, instead of doing one fixed search. Graph RAG builds a knowledge graph of entities and relationships from the documents and retrieves over that graph, which helps with questions that need connections across many documents, such as "what are the main themes in this corpus".
Also asked as: what is agentic rag · how does agentic rag work · agentic rag vs graph rag · what is graph rag · how to implement graph rag · agentic rag vs rag · multi-agent rag
Microsoft's Graph RAG paper describes building an entity graph with community summaries and shows gains on global, sense-making questions over a corpus [13]. HyDE, another retrieval upgrade, has the model write a hypothetical answer first and embeds that to search, which helps when questions are short and vague [11]. My advice to students is to climb the levels above in order. Agentic loops and graphs are level-eight problems; most systems I audit are stuck at level two.
Is RAG still relevant with long context windows?
Yes. Long context windows let you paste more into a prompt, but they do not tell you which of a million documents to paste, they cost per token on every question, and models still attend unevenly across long inputs. Retrieval is how you choose what deserves the context window. Longer windows make RAG easier, because you can send more and better passages, not unnecessary.
Also asked as: is rag dead · rag vs long context · do we still need rag · will long context replace rag · future of rag
Liu et al.'s finding that accuracy drops when relevant information sits in the middle of a long context is the empirical reason retrieval and reranking still matter [8]. The economics are the practical reason: a 200-page manual is hundreds of thousands of tokens; five reranked passages are a few thousand.
What are common RAG interview questions?
Expect to explain the two phases, to compare RAG with fine-tuning, to describe chunking trade-offs, to justify hybrid search and reranking, to define recall and faithfulness, and to walk through a failure you debugged. Interviewers are testing whether you have operated a RAG system, not whether you can recite the paper.
Also asked as: rag interview questions · rag interview questions and answers · genai interview questions rag · llm interview questions retrieval
The questions that separate candidates, from the interviewing I have done for SCULT INDIA and from the sessions I run for students:
- How do you know your retrieval is working before you look at the answers?
- A user says the bot gave a wrong answer. Walk me through your debugging, in order.
- Why would you add BM25 to a system that already has embeddings?
- What happens when two retrieved passages disagree?
- How do you stop the model answering when the documents do not contain the answer?
If you can answer those five from experience, you can build the system. The rest of this guide is the experience, compressed.
Where can I learn RAG properly?
Read the original paper and the 2023 survey for the map, build one pipeline by hand without a framework so you understand every stage, then evaluate it on your own questions. That order, paper, build, measure, teaches more than any course, and it is exactly what I ask of students who want to work on retrieval systems with me.
Also asked as: rag course · best rag course · rag roadmap · rag tutorial · learn rag from scratch · rag for beginners · rag project ideas
A good first project is a RAG system over your own college's notes, syllabus and previous-year papers, because you can judge the answers yourself and the documents are messy in exactly the ways real client documents are. I mentor students through TechVerse Enclave on projects like this, and the ones who finish have something no certificate gives them: a system they can explain failure by failure. If you want to talk RAG at your college or hackathon, my email is pranjulrathour41@gmail.com and the invite page is at pranjulrathour.scult.in/invite.
Sources
- Lewis et al., Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks (NeurIPS 2020)arxiv.org
- Gao et al., Retrieval-Augmented Generation for Large Language Models: A Survey (2023)arxiv.org
- Karpukhin et al., Dense Passage Retrieval for Open-Domain Question Answering (2020)arxiv.org
- Robertson & Zaragoza, The Probabilistic Relevance Framework: BM25 and Beyond (2009)staff.city.ac.uk
- Cormack, Clarke & Buettcher, Reciprocal Rank Fusion outperforms Condorcet and individual rank learning methods (SIGIR 2009)plg.uwaterloo.ca
- Nogueira & Cho, Passage Re-ranking with BERT (2019)arxiv.org
- Johnson, Douze & Jégou, Billion-scale similarity search with GPUs (FAISS, 2017)arxiv.org
- Liu et al., Lost in the Middle: How Language Models Use Long Contexts (2023)arxiv.org
- Es et al., RAGAS: Automated Evaluation of Retrieval Augmented Generation (2023)arxiv.org
- Barnett et al., Seven Failure Points When Engineering a Retrieval Augmented Generation System (2024)arxiv.org
- Gao et al., Precise Zero-Shot Dense Retrieval without Relevance Labels (HyDE, 2022)arxiv.org
- Asai et al., Self-RAG: Learning to Retrieve, Generate, and Critique through Self-Reflection (2023)arxiv.org
- Edge et al., From Local to Global: A Graph RAG Approach to Query-Focused Summarization (Microsoft, 2024)arxiv.org
- MTEB: Massive Text Embedding Benchmark leaderboard, Hugging Facehuggingface.co
- pgvector: open-source vector similarity search for Postgresgithub.com
- Qdrant documentationqdrant.tech
- RAG.NextUpgrad source code, Pranjul Rathourgithub.com
- Why my RAG platform says 'I don't know', Pranjul Rathourpranjulrathour.scult.in




