What is a context window, and what is context engineering? Token limits, why bigger isn't always better, and deciding what a model actually sees

A context window is the maximum amount of text, in tokens, a model can hold in one call: the prompt, retrieved documents, conversation history and the answer all share this one budget. What happens when you exceed it, why a huge context window doesn't remove the need for retrieval, the lost-in-the-middle effect, and context engineering, the discipline of deciding exactly what enters that budget and in what order.

Pranjul Rathour
Pranjul Rathour

Key takeaways

  • The context window is a shared token budget across the system prompt, retrieved documents, conversation history and the model's own answer.
  • A bigger context window changes what fits, not how well a model uses everything inside it; the lost-in-the-middle effect is real and measured.
  • Context engineering is the practice of deciding what enters the window and in what order, distinct from prompt engineering, which is about phrasing.
  • Long context does not remove the case for retrieval; it changes the trade-off between retrieval precision and raw context cost and latency.
  • Track token usage in production the way you track any other resource: log it, budget it, and alert when a request approaches the limit.

What is a context window?

A context window is the maximum number of tokens a model can process in a single call, shared across everything that goes in: the system prompt, retrieved documents, the conversation history, and the model's own output. If a request's total token count exceeds the window, either the call fails outright or the oldest content is silently truncated, depending on how the application handles it, which is why token counting is something a serious application tracks explicitly rather than discovers by accident [2][7].

Also asked as: what is a context window · context window meaning · what does context window mean in ai · llm context window explained · how big is a context window · context window vs token limit · what happens if you exceed context window

RAG.NextUpgrad, the platform I run, budgets this explicitly: a fixed slice for the system prompt, a capped slice for retrieved passages, and whatever remains for conversation history, because letting any one of those grow unchecked starves the others.

A context window is not a bigger brain. It is a bigger desk. What matters is still what you put on it and where. Pranjul Rathour

Does a bigger context window mean the model understands more?

Not proportionally. Models with very large context windows can accept far more tokens, but studies on how models actually use long contexts found a real "lost in the middle" effect: performance on information placed in the middle of a long context is measurably worse than the same information placed near the beginning or end [1]. A larger window changes what fits into a call; it does not guarantee the model weighs every part of that content equally well, which is why "needle in a haystack" tests exist specifically to measure this per model and context length [5].

Also asked as: does bigger context window mean better · lost in the middle effect · context window size vs performance · does more context help llm accuracy · needle in a haystack test · long context llm limitations

Does a huge context window remove the need for RAG?

No. Even with a context window large enough to hold an entire document collection, sending all of it on every call is slower and more expensive than retrieving the few passages actually relevant to a question, and the lost-in-the-middle effect means precision still matters: a well-retrieved, well-ordered handful of passages regularly outperforms a huge unfiltered dump of text on the same question [1][8]. Long context changes the trade-off, sometimes it is cheaper to skip retrieval for a single small document, but it does not eliminate the case for retrieval at real scale.

Also asked as: does long context replace rag · is rag still needed with long context models · long context vs retrieval augmented generation · when to use long context instead of rag · rag vs stuffing the context window

My longer notes on why RAG still matters, and on chunking specifically, are on the RAG explainer and the chunking page [8][9].

Pranjul Rathour
Pranjul Rathour

What is context engineering?

Context engineering is the discipline of deciding exactly what enters a model's context window, in what order, and in what form, as a distinct skill from prompt engineering, which is about how you phrase instructions [6]. It covers what system instructions to include, which retrieved passages to keep and how to order them, how much conversation history to carry forward and how to summarise the rest, and what tool definitions or examples earn a place in an already-constrained budget. Anthropic's framing treats it as engineering a shared, finite resource rather than writing a clever paragraph [6].

Also asked as: what is context engineering · context engineering vs prompt engineering · context engineering explained · how to do context engineering · context engineering for agents · what does a context engineer do

How do I manage a context budget in a real application?

Fix an explicit token budget per section before you write a line of code: so many tokens for the system prompt, a capped number for retrieved passages, so many for recent conversation turns, with the rest reserved for the model's own output. Summarise or drop older conversation history rather than letting it grow unbounded. Order retrieved passages by relevance with the most important ones near the start or end, not buried in the middle, given the lost-in-the-middle effect [1]. Count tokens with the provider's actual tokenizer, not a rough word-count estimate, since the two diverge enough to matter near a limit [7].

Also asked as: how to manage context window budget · token budget for llm app · context window management strategies · how to truncate conversation history · llm token counting · managing long conversations with llm

What happens when I exceed the context window?

Depending on the provider and how your application calls it, the request either fails with an explicit error, or, in poorly built applications, silently truncates the oldest content, which can quietly drop the system prompt or the earliest turns of a conversation without any visible error. Neither is acceptable in production without you choosing it deliberately. Handle this explicitly: check token counts before sending, and truncate or summarise on your own terms rather than letting the provider or a library do it silently.

Also asked as: what happens if you exceed token limit · context window exceeded error · llm truncation behavior · how to handle context length exceeded · token limit error fix

How much does context length affect cost and latency?

Directly and significantly: nearly every provider bills by total tokens processed, input and output, so a request that sends more context costs more regardless of how much of it the model actually needed, and larger inputs generally take longer to process even before generation begins [3][10]. This is the practical argument for retrieval and careful context budgets even when a model's window could technically hold everything: cost and latency scale with what you send, not with what would have been sufficient.

Also asked as: context length and cost · does more context cost more · llm latency and context size · context window pricing · how context size affects api cost

Context window and context engineering interview questions

Explain what a context window is and what shares its budget. Describe the lost-in-the-middle effect and its practical implication for ordering retrieved content. Explain why a long-context model still benefits from retrieval. Define context engineering and contrast it with prompt engineering. Describe how you would budget tokens across a RAG conversation. The strongest answer names a specific ordering or truncation decision you made and why.

Also asked as: context window interview questions · context engineering interview · llm token limit interview questions · rag context interview questions

Where should I start?

Count the tokens in your current system prompt, your typical retrieved passages, and a typical conversation history, and see how much of the window each actually uses. Most budgets are more lopsided than people assume. For a hands-on session on context and prompt engineering for real applications, email pranjulrathour41@gmail.com or use pranjulrathour.scult.in/invite.

Sources

  1. Liu et al., Lost in the Middle: How Language Models Use Long Contexts (2023)arxiv.org
  2. Anthropic, context windows documentationdocs.anthropic.com
  3. OpenAI, managing context lengthplatform.openai.com
  4. Google, long context in Gemini modelsai.google.dev
  5. Kamradt, Needle In A Haystack, long-context evaluationgithub.com
  6. Anthropic, effective context engineering for AI agentsanthropic.com
  7. Anthropic, tokenizer and token countingdocs.anthropic.com
  8. What is RAG in AI, Pranjul Rathourpranjulrathour.github.io
  9. What is chunking in RAG, Pranjul Rathourpranjulrathour.github.io
  10. How much does an LLM API cost, Pranjul Rathourpranjulrathour.github.io
  11. What is agentic RAG and Graph RAG, Pranjul Rathourpranjulrathour.github.io
Pranjul Rathour
Pranjul Rathour
GenAI Engineer · Kanpur, Uttar Pradesh, India

GenAI engineer and AI product builder with 2+ years shipping production-grade AI systems: RAG pipelines, fine-tuned LLMs, hybrid retrieval and multi-modal apps across vision, speech and OCR, architected end to end from ingestion to deployment. Leads engineering for SCULT INDIA's 14-member team, founded the 500+ member TechVerse Enclave community and has mentored 200+ students. Three hackathon first prizes: Changethon 2025 (IIT Roorkee), Product Genesis at Vividhotsava 2025 (CSJMU Kanpur) and BYTEBATTLE (MeetKats).

Open to GenAI roles, hackathon judging, mentorship sessions and guest talks at colleges. Email pranjulrathour41@gmail.com.

Keep reading

More on RAG & retrieval

All RAG & retrieval guides →
Campus talks · hackathon judging · mentoring

Want this as a live session at your college?

Open to GenAI roles, hackathon judging, mentorship sessions and guest talks at colleges.

Pranjul Rathour, GenAI engineer in Kanpur