RAG.NextUpgrad, the architecture

Most RAG tutorials end at 'and then we call the LLM'. Production starts there.

RAG.NextUpgrad, the architecture, slide 1 of 9
Slides
01 / 09
RAG.NextUpgrad, the architecture, slide 1 of 9
01
RAG.NextUpgrad, the architecture, slide 2 of 9
02
RAG.NextUpgrad, the architecture, slide 3 of 9
03
RAG.NextUpgrad, the architecture, slide 4 of 9
04
RAG.NextUpgrad, the architecture, slide 5 of 9
05
RAG.NextUpgrad, the architecture, slide 6 of 9
06
RAG.NextUpgrad, the architecture, slide 7 of 9
07
RAG.NextUpgrad, the architecture, slide 8 of 9
08
RAG.NextUpgrad, the architecture, slide 9 of 9
09
The caption

Most RAG tutorials end at 'and then we call the LLM'. Production starts there. RAG.NextUpgrad is the platform I built for real users: hybrid dense + BM25 retrieval fused with Reciprocal Rank Fusion, re-ranking, a confidence gate, streaming citations — and the unglamorous parts: three vector stores behind one interface, LLM and embedding providers that fail over automatically, prompt-injection protection on retrieved text, and ~220 MB of RAM so it fits a free tier. Here's the architecture, endpoint by endpoint. Built at NextUpgrad Web Solutions; code on GitHub. I also build tools.scult.in — 15 free tools, 1,200+ prompts and a 50,000-skill library, free for anyone, no signup.

Keep flipping

More carousels

All 59 decks →
Campus talks · hackathon judging · mentoring

Want this as a live session at your college?

Open to GenAI roles, hackathon judging, mentorship sessions and guest talks at colleges.

Pranjul Rathour, GenAI engineer in Kanpur