How to evaluate a RAG system: recall, faithfulness, RAGAS, building a test set and the metrics that actually predict user complaints

Evaluate retrieval and generation separately. How to build a labelled question set from real usage, measure recall at k and MRR for retrieval, faithfulness and answer relevance for generation, use RAGAS and LLM judges without fooling yourself, run evaluation in CI, and read the numbers that predict whether users will trust the system. From a production RAG platform with 196 tests.

Walking a room through evaluation criteria
Walking a room through evaluation criteria

Key takeaways

  • Evaluate retrieval first and separately: if the right passage is not in the top k, no prompt or model can fix the answer.
  • Twenty labelled real questions beat a thousand synthetic ones for finding what is broken. Grow the set from production failures.
  • Faithfulness, whether every claim is supported by the retrieved passages, is the generation metric that maps to user trust.
  • LLM-as-judge metrics such as RAGAS are useful and cheap, and they need a human-labelled sample to calibrate against.
  • Put the evaluation in CI. A RAG system that is not re-evaluated on every prompt, chunking or model change drifts silently.

How do I evaluate a RAG system?

In two separate stages. Retrieval: build a set of real questions each labelled with the passage that answers it, run the retriever, and measure how often the right passage appears in the top k results. Generation: for each question, judge whether the answer is faithful to the retrieved passages, whether it actually answers the question, and whether it cites correctly. Then track both sets of numbers on every change, in CI, so a chunking tweak or a model upgrade cannot silently make the system worse.

Also asked as: how to evaluate rag · rag evaluation · rag evaluation metrics · how to evaluate rag pipeline · how to evaluate rag chatbot · how to test rag system · rag evaluation framework · rag benchmark · evaluating retrieval augmented generation · rag quality metrics

RAG.NextUpgrad, the production platform I built in 2026, carries 196 automated tests, and the ones that catch real regressions are the retrieval recall tests over a labelled question set [10]. Everything on this page is what those tests taught me, with the papers underneath. My shorter notes on the metrics are on the portfolio [11][12].

If the right passage is not in the top thirty, stop. No prompt, no model, no reranker can answer from a document it never saw. Measure that first, and measure it alone. Pranjul Rathour, from building RAG.NextUpgrad

Why evaluate retrieval and generation separately?

Because they fail differently and are fixed differently. A wrong answer can come from retrieval returning the wrong chunk, which is a chunking or search problem, or from the model ignoring a correct chunk, which is a prompt or model problem. A single end-to-end score hides which one happened. Measure retrieval with the labelled passages, measure generation given the passages actually retrieved, and the failure points name themselves.

Also asked as: retrieval vs generation evaluation · why evaluate retrieval separately · rag component evaluation · end to end rag evaluation · rag failure analysis · rag debugging

Barnett et al.'s seven failure points map almost one to one onto the categories that fall out of this loop [7].

How do I build an evaluation set for RAG?

Start with twenty real questions: from logs if you have users, from the people who own the documents if you do not, or from yourself reading the documents and writing what a user would ask. For each, mark the chunk or passage that answers it, and write the answer you would accept. Include hard cases: questions with no answer in the corpus, questions whose answer spans two sections, and questions using different words from the document. Grow the set from every production complaint. Twenty is enough to start; two hundred is enough to trust.

Also asked as: rag test set · how to create evaluation dataset for rag · rag golden dataset · rag ground truth · synthetic test set for rag · rag evaluation dataset · how many questions to evaluate rag · rag eval examples

Synthetic questions generated from chunks are useful to widen coverage cheaply, and ARES shows how to use them with a small human-labelled set for calibration [5], but they tend to use the document's own phrasing, which overstates recall. Keep the real ones as the reference.

What are the retrieval metrics?

Recall at k: the share of questions for which a correct passage appears in the top k results; the metric that matters most, at the k you actually send to the reranker. Precision at k: the share of the top k that are relevant, which matters for prompt cost. Mean reciprocal rank: how high the first correct passage ranks, on average. Hit rate is recall at k under another name. Report recall at your retrieval k and at your generation k, because they answer different questions.

Also asked as: recall at k rag · rag recall · precision at k · mrr retrieval · hit rate rag · retrieval metrics · ndcg rag · context recall · context precision · how to measure retrieval quality

The definitions are classical information retrieval, covered in the Manning textbook's evaluation chapter [8]; BEIR is the benchmark most embedding models are compared on [6], and it is not your corpus, so measure your own.

What are the generation metrics?

Faithfulness: the fraction of claims in the answer that are supported by the retrieved passages, which is the metric for hallucination. Answer relevance: whether the answer addresses the question asked. Context relevance or precision: whether the passages sent were needed. Citation correctness: whether each citation points to a passage that supports the sentence. Refusal accuracy: whether the system said "not in the documents" exactly when it should. RAGAS defines the first three and scores them with an LLM judge [1][2].

Also asked as: faithfulness rag · answer relevance rag · ragas metrics · rag hallucination metric · groundedness score · context relevance · citation accuracy rag · answer correctness · rag generation metrics · llm as a judge rag

Can I trust LLM-as-a-judge metrics?

Partly. Judge models correlate with human ratings well enough to be useful for tracking changes, and they are cheap enough to run on every commit; Zheng et al. measured the agreement and its biases, including position and verbosity effects [3]. They are not a substitute for a human-labelled sample, because a judge can be faithful to the wrong passage or fooled by confident wording. Calibrate: label fifty answers by hand, compare with the judge, and only then trust the judge on the other thousand.

Also asked as: llm as a judge reliability · llm as a judge bias · ragas accuracy · can you trust ragas · automated rag evaluation vs human · llm judge calibration · g-eval · rag evaluation without ground truth

Walking a room through evaluation criteria
Walking a room through evaluation criteria

How do I evaluate hybrid search and reranking?

Ablate. Run recall at k for dense retrieval alone, BM25 alone, and the fused list, on the same question set. Then run recall over the top five after reranking versus the top five before. Each stage should raise recall at its output k, and if a stage does not, remove it. In RAG.NextUpgrad the fused list beat either retriever alone on the labelled set, and the reranker's job was visible as recall at five rising while recall at thirty stayed flat [10].

Also asked as: evaluate hybrid search · evaluate reranker · ablation study rag · bm25 vs dense evaluation · reranking evaluation metrics · how to know if reranker helps · rag ablation

How do I evaluate chunking and embedding model choices?

The same way: fix the question set, vary one component, compare recall at k. Index the corpus at three chunk sizes and compare. Embed with two models and compare. Prepend headings and compare. Every configuration decision in a RAG system is an experiment against the labelled set, and the set is what makes the decision defensible rather than a guess copied from a tutorial. Models that look worse on a public leaderboard sometimes win on your documents.

Also asked as: evaluate chunking strategy · compare embedding models for rag · chunk size evaluation · embedding model evaluation rag · rag configuration testing · rag experiments · rag a b testing

How do I run RAG evaluation in CI?

Store the question set and labels in the repository. Write a test that indexes a fixed fixture corpus, runs retrieval for every question, and asserts recall at k above a threshold you set from the current baseline. Write a second test that generates answers for a subset and asserts faithfulness and refusal accuracy above thresholds, using a cheap judge model, and fails the build on regression. Cache embeddings so the test runs in minutes. Review the failures as part of code review, not after deployment.

Also asked as: rag testing in ci · automated rag testing · regression testing for rag · rag unit tests · test rag pipeline pytest · continuous evaluation llm · rag evaluation automation

def test_retrieval_recall_at_10(index, eval_set):
    hits = 0
    for q in eval_set:
        top = index.search(q.question, k=10)
        hits += any(p.id in q.relevant_ids for p in top)
    recall = hits / len(eval_set)
    assert recall >= 0.85, f"recall@10 dropped to {recall:.2f}"   # baseline set from the last release

That test, and its siblings for BM25, fusion and reranking, is a large part of why RAG.NextUpgrad has 196 tests and why chunking changes ship with confidence [10].

What numbers should I aim for?

Recall at the reranker's input k, say thirty, above ninety percent on your labelled set, or retrieval is the problem. Recall at the generation k, say five, above eighty-five. Faithfulness above ninety-five percent on questions with an answer, and refusal accuracy near one hundred on questions without one. These are targets from production experience, not universal thresholds; the right numbers depend on how costly a wrong answer is for your users.

Also asked as: good recall for rag · rag benchmark scores · what is a good faithfulness score · rag accuracy target · rag performance benchmark · acceptable rag metrics · rag kpi

What do users actually complain about, and which metric predicts it?

Confident wrong answers, which faithfulness predicts. "It said it does not know" when the answer was there, which recall predicts. Slow answers, which no quality metric predicts, so track latency alongside. Answers that ignore the newest document, which is index freshness, so track it. The metrics above cover the first two; the last two are operations, and they belong on the same dashboard.

Also asked as: rag user complaints · rag monitoring · rag observability · rag latency · rag freshness · production rag metrics · rag dashboard · what to log in rag

RAG evaluation interview questions

Why evaluate retrieval separately from generation? Define recall at k and faithfulness. How would you build an evaluation set with no users yet? How do you know your reranker helps? What are the risks of LLM-as-a-judge? How would you put evaluation in CI? What number would you look at first when a user reports a wrong answer? Every answer is above, and a repository with a labelled set and a recall test is the strongest evidence you can bring.

Also asked as: rag evaluation interview questions · llm evaluation interview · how to evaluate rag interview answer · ml evaluation interview questions genai

Where should I start?

Write twenty real questions against a document set you know, mark the passage that answers each, and compute recall at ten for your current retriever. Read the misses. That afternoon will tell you more about your system than any framework's dashboard. For a hands-on session on building and evaluating retrieval systems at your college or team, email pranjulrathour41@gmail.com or use pranjulrathour.scult.in/invite.

Sources

  1. Es et al., RAGAS: Automated Evaluation of Retrieval Augmented Generation (2023)arxiv.org
  2. RAGAS documentationdocs.ragas.io
  3. Zheng et al., Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena (2023)arxiv.org
  4. Chen et al., Benchmarking Large Language Models in Retrieval-Augmented Generation (2023)arxiv.org
  5. Saad-Falcon et al., ARES: An Automated Evaluation Framework for Retrieval-Augmented Generation Systems (2023)arxiv.org
  6. Thakur et al., BEIR: A Heterogeneous Benchmark for Zero-shot Evaluation of Information Retrieval Models (2021)arxiv.org
  7. Barnett et al., Seven Failure Points When Engineering a RAG System (2024)arxiv.org
  8. Manning, Raghavan & Schütze, Introduction to Information Retrieval, chapter 8: Evaluationnlp.stanford.edu
  9. Liu et al., Lost in the Middle (2023)arxiv.org
  10. RAG.NextUpgrad source code, 196 tests, Pranjul Rathourgithub.com
  11. How to evaluate a RAG system: recall, faithfulness and the questions that matter, Pranjul Rathourpranjulrathour.scult.in
  12. Evaluating retrieval quality without labelled data, Pranjul Rathourpranjulrathour.scult.in
Pranjul Rathour
Pranjul Rathour
GenAI Engineer · Kanpur, Uttar Pradesh, India

GenAI engineer and AI product builder with 2+ years shipping production-grade AI systems: RAG pipelines, fine-tuned LLMs, hybrid retrieval and multi-modal apps across vision, speech and OCR, architected end to end from ingestion to deployment. Leads engineering for SCULT INDIA's 14-member team, founded the 500+ member TechVerse Enclave community and has mentored 200+ students. Three hackathon first prizes: Changethon 2025 (IIT Roorkee), Product Genesis at Vividhotsava 2025 (CSJMU Kanpur) and BYTEBATTLE (MeetKats).

Open to GenAI roles, hackathon judging, mentorship sessions and guest talks at colleges. Email pranjulrathour41@gmail.com.

Keep reading

More on RAG & retrieval

All RAG & retrieval guides →
Campus talks · hackathon judging · mentoring

Want this as a live session at your college?

Open to GenAI roles, hackathon judging, mentorship sessions and guest talks at colleges.

Pranjul Rathour, GenAI engineer in Kanpur