What is agentic RAG and Graph RAG? Query rewriting, multi-step retrieval, knowledge graphs, and when a plain RAG pipeline is not enough

Agentic RAG puts a planning loop around retrieval: rewrite the query, choose among retrievers, check results, retrieve again. Graph RAG builds a knowledge graph over documents and retrieves through entities and communities. What each adds, what each costs, HyDE, Self-RAG, RAPTOR and corrective patterns, and how to decide whether your system needs any of it. From a production RAG platform that mostly does not.

Presenting to a room
Presenting to a room

Key takeaways

  • Agentic RAG adds decisions to the pipeline: whether to retrieve, how to rewrite, which source to use, whether the results are good enough to answer.
  • Graph RAG adds structure: entities, relationships and community summaries, which help with global questions about a corpus that chunk retrieval cannot answer.
  • Both multiply cost and latency. Most systems that struggle are failing at chunking, hybrid search or reranking, not at planning.
  • Query rewriting, HyDE and a retrieval-check step are the cheap agentic upgrades; a full multi-agent loop is rarely justified.
  • Measure recall and faithfulness before and after. If the plain pipeline is above 90 percent recall, agentic layers will not pay for themselves.

What is agentic RAG?

Agentic RAG wraps the retrieval pipeline in a decision loop run by the model: decide whether retrieval is needed at all, rewrite or decompose the query, choose which index or tool to search, inspect what came back, and retrieve again or differently if the evidence is weak, before generating. A plain RAG pipeline does one fixed retrieve-then-generate pass [1]; an agentic one lets the model steer, using patterns from ReAct and its descendants [9]. It answers harder, multi-step questions at the cost of more model calls, more latency and more ways to fail.

Also asked as: what is agentic rag · agentic rag explained · agentic rag vs rag · agentic rag vs traditional rag · how does agentic rag work · agentic rag architecture · agentic retrieval · rag agents · what is agentic retrieval augmented generation · multi agent rag

RAG.NextUpgrad, the production platform I built, is deliberately not agentic in the full sense: it has one pass with hybrid retrieval, reranking and a confidence gate, and that pass scores above ninety percent recall on the labelled set [15]. This page is about what the agentic and graph layers add, what they cost, and the test I use to decide whether a system needs them. My notes on multi-hop questions and query rewriting are on the portfolio [13][14].

Every agentic RAG system I have audited that was struggling had a chunking or ranking problem underneath. Fix the pipeline first. Add planning when the pipeline is good and the questions still fail. Pranjul Rathour

How does agentic RAG differ from a plain RAG pipeline?

A plain pipeline is a fixed sequence: embed, retrieve, rerank, generate, once. An agentic system inserts model decisions between the stages: a router that picks the retriever, a rewriter that improves the query, a grader that checks whether retrieved passages actually answer the question, and a loop that retries with a different query or source. The decisions are prompts or small classifiers; the loop is a graph with a step limit. Every decision is a place for the model to be right or wrong, and a place to measure.

Also asked as: agentic rag vs naive rag · advanced rag vs agentic rag · rag pipeline vs rag agent · difference between rag and agentic rag · agentic rag components · agentic rag workflow · rag router · rag grader

What is query rewriting, and what is HyDE?

Query rewriting has the model turn the user's question into a better retrieval query: expanding abbreviations, adding likely terms, splitting a compound question into sub-questions, or removing conversational noise. HyDE, hypothetical document embeddings, goes further: the model writes a short hypothetical answer, and that text is embedded and used to search, because an answer-shaped passage is closer in vector space to real answer passages than a short question is [5]. Both are one extra model call and often the biggest agentic-style gain available.

Also asked as: query rewriting rag · what is hyde · hypothetical document embeddings · query expansion rag · query decomposition rag · multi query rag · rag fusion · query transformation rag · how to improve rag queries · step back prompting rag

Query rewriting is the one agentic element I add to almost every system, because conversational follow-ups such as "and for the other plan?" are unretrievable as written [14].

What are Self-RAG and corrective RAG?

Self-RAG trains a model to emit reflection tokens that decide whether to retrieve, whether a passage is relevant, whether the generation is supported, and whether it is useful, so retrieval and self-critique happen inside generation [3]. Corrective RAG adds a lightweight evaluator that grades retrieved documents as correct, incorrect or ambiguous and triggers web search or query refinement when retrieval looks wrong [4]. Both formalise the same idea a confidence gate implements more simply: check the evidence before you trust it.

Also asked as: self rag · self reflective rag · corrective rag · crag rag · rag with self critique · rag grading retrieved documents · adaptive rag · rag evaluation loop · reflection tokens

The confidence gate in RAG.NextUpgrad is the minimal version of this family: it inspects reranker scores and passage agreement and refuses when support is weak, with no extra model call [15]. Start there; graduate to graded loops when the refusal rate on answerable questions is measurably too high.

What is Graph RAG?

Graph RAG builds a knowledge graph from the corpus, extracting entities and the relationships between them with a model, clusters the graph into communities, writes a summary for each community, and then retrieves through the graph: by entity for specific questions, by community summary for global ones. Microsoft's Graph RAG paper showed gains on query-focused summarisation, questions such as "what are the main themes across these documents", where chunk retrieval has nothing to retrieve because the answer is not in any single chunk [6][7].

Also asked as: what is graph rag · graph rag explained · graphrag microsoft · knowledge graph rag · graph rag vs vector rag · graph rag vs rag · how does graph rag work · graph rag architecture · when to use graph rag · graph database for rag

RAPTOR reaches a similar goal, multi-level retrieval, by recursively clustering and summarising chunks into a tree, without extracting an explicit graph [8]. Both trade index-time cost for the ability to answer questions about the whole.

When do I actually need agentic RAG or Graph RAG?

When the plain pipeline is measurably good and specific question types still fail. Multi-hop questions that need two retrievals, "who leads the team that built the tool mentioned in the June report", need interleaved retrieval and reasoning, which IRCoT and agentic loops provide [10]. Global questions about a corpus need Graph RAG or RAPTOR. Questions spanning several systems need a router. Everything else, which is most questions in most products, needs chunking, hybrid search and reranking done well.

Also asked as: when to use agentic rag · do i need agentic rag · when is graph rag worth it · agentic rag use cases · graph rag use cases · multi hop rag · complex questions rag · rag for multiple data sources · rag limitations agentic solution

Presenting to a room
Presenting to a room

How do I build an agentic RAG system?

Model it as a graph with explicit state: nodes for rewrite, retrieve, grade, generate and refuse, edges chosen by the grader's output, a step limit and a token budget. LangGraph is the common framework for this shape [11]; the same loop is a couple of hundred lines without one. Log every decision. Evaluate the agentic system against the plain pipeline on the same labelled set and keep it only where it measurably wins. Anthropic's guidance applies: the simplest architecture that works, and the model's judgement only where it is actually needed [12].

Also asked as: how to build agentic rag · agentic rag langgraph · agentic rag tutorial · agentic rag with langchain · agentic rag llamaindex · agentic rag python · agentic rag implementation · rag agent framework · build rag agent

What does agentic RAG cost?

Several model calls per question instead of one, plus latency that users notice, plus a larger failure surface: a rewriter that drifts, a grader that is too strict, a loop that runs to its limit. A rewriter and a grader roughly double cost; a full multi-agent design can multiply it several times. Graph RAG shifts cost to index time, with many model calls to extract entities and write summaries, which repeats when documents change. Budget both before building, and measure whether the gain in recall or answer quality pays for it.

Also asked as: agentic rag cost · agentic rag latency · graph rag cost · is graph rag expensive · agentic rag performance · agentic rag trade offs · graph rag indexing cost · rag cost vs quality

What are common mistakes with agentic and Graph RAG?

Adding an agent loop to a pipeline with bad chunking. No step limit, so a hard question loops until the budget dies. A grader that rejects good passages and triggers endless retries. Graph RAG for a corpus of precise procedures where chunk citations were the requirement. Not logging decisions, so failures cannot be traced. Evaluating the agentic system alone instead of against the plain baseline. Every one of these is avoided by measuring first and adding layers only where the measurement demands them.

Also asked as: agentic rag mistakes · agentic rag failure modes · graph rag limitations · problems with agentic rag · agentic rag debugging · when agentic rag fails · graph rag drawbacks

Agentic RAG interview questions

Define agentic RAG and Graph RAG and say what each fixes. Explain query rewriting and HyDE. Describe Self-RAG or corrective RAG in a sentence. Give a question a plain pipeline cannot answer and say which layer fixes it. Explain how you would decide whether to add an agentic layer, and how you would measure the result. Describe the cost trade-offs. The strongest answer includes a pipeline you measured and a layer you chose not to add.

Also asked as: agentic rag interview questions · graph rag interview questions · advanced rag interview questions · rag architecture interview · multi hop rag interview

Where should I start?

Measure your plain pipeline's recall on twenty labelled questions. Read the misses. If they are conversational follow-ups, add query rewriting. If they are multi-hop, add a retrieval check and one retry. If they are global questions about the corpus, prototype RAPTOR or Graph RAG on a subset and compare. Add nothing you cannot measure. For a hands-on session on retrieval systems from the notebook to the operator level, email pranjulrathour41@gmail.com or use pranjulrathour.scult.in/invite.

Sources

  1. Lewis et al., Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks (2020)arxiv.org
  2. Gao et al., Retrieval-Augmented Generation for Large Language Models: A Survey (2023)arxiv.org
  3. Asai et al., Self-RAG: Learning to Retrieve, Generate, and Critique through Self-Reflection (2023)arxiv.org
  4. Yan et al., Corrective Retrieval Augmented Generation (CRAG, 2024)arxiv.org
  5. Gao et al., Precise Zero-Shot Dense Retrieval without Relevance Labels (HyDE, 2022)arxiv.org
  6. Edge et al., From Local to Global: A Graph RAG Approach to Query-Focused Summarization (Microsoft, 2024)arxiv.org
  7. Microsoft GraphRAG documentationmicrosoft.github.io
  8. Sarthi et al., RAPTOR: Recursive Abstractive Processing for Tree-Organized Retrieval (2024)arxiv.org
  9. Yao et al., ReAct: Synergizing Reasoning and Acting in Language Models (2022)arxiv.org
  10. Trivedi et al., Interleaving Retrieval with Chain-of-Thought Reasoning for Knowledge-Intensive Multi-Step Questions (IRCoT, 2022)arxiv.org
  11. LangGraph documentationlangchain-ai.github.io
  12. Anthropic, Building effective agents (2024)anthropic.com
  13. Multi-hop questions in RAG systems, Pranjul Rathourpranjulrathour.scult.in
  14. Query rewriting and HyDE for retrieval, Pranjul Rathourpranjulrathour.scult.in
  15. RAG.NextUpgrad source code, Pranjul Rathourgithub.com
Pranjul Rathour
Pranjul Rathour
GenAI Engineer · Kanpur, Uttar Pradesh, India

GenAI engineer and AI product builder with 2+ years shipping production-grade AI systems: RAG pipelines, fine-tuned LLMs, hybrid retrieval and multi-modal apps across vision, speech and OCR, architected end to end from ingestion to deployment. Leads engineering for SCULT INDIA's 14-member team, founded the 500+ member TechVerse Enclave community and has mentored 200+ students. Three hackathon first prizes: Changethon 2025 (IIT Roorkee), Product Genesis at Vividhotsava 2025 (CSJMU Kanpur) and BYTEBATTLE (MeetKats).

Open to GenAI roles, hackathon judging, mentorship sessions and guest talks at colleges. Email pranjulrathour41@gmail.com.

Keep reading

More on RAG & retrieval

All RAG & retrieval guides →
Campus talks · hackathon judging · mentoring

Want this as a live session at your college?

Open to GenAI roles, hackathon judging, mentorship sessions and guest talks at colleges.

Pranjul Rathour, GenAI engineer in Kanpur