What is an LLM? Large language models explained: tokens, context windows, parameters, open vs closed models, and what they cost to run

A large language model predicts the next token, and everything else follows from that. How LLMs are trained, what parameters, tokens and context windows mean, how GPT, Claude, Gemini, Llama and Mistral differ, what hallucination is, how to run one locally, and how to reason about cost. Written from shipping LLM products.

On the mic
On the mic

Key takeaways

  • An LLM is a neural network trained to predict the next token over enormous text; chat, code and reasoning emerge from that objective plus instruction tuning.
  • Parameters are the model's size, tokens are what it reads and writes, the context window is how much it can read at once. Each has a cost.
  • Hallucination is not a bug to patch but a property of prediction. Retrieval, grounding and evaluation are how you manage it.
  • Open-weight models (Llama, Mistral, Gemma, Qwen) run on your hardware; closed models (GPT, Claude, Gemini) run behind APIs. Choose by control, cost and quality for your task.
  • Cost is tokens in plus tokens out, times the price per million, times requests. Measure it on real traffic before you promise a budget.

What is an LLM?

A large language model is a neural network, almost always a transformer, trained on a very large body of text to predict the next token, a word or word-piece, given everything before it. That one objective, repeated over trillions of tokens, produces a model that can complete, summarise, translate, answer and write code. Chat behaviour is added afterwards by instruction tuning and preference training. "Large" refers to the parameter count, from about a billion to over a trillion.

Also asked as: owasp top 10 llm · what does llm stand for in ai · best llm for coding · how llm works · best local llm for coding · best llm visibility tracking software · how to fix llm memory evaluation bias with readerfacing artifacts · is there a way to interrupt text generation in an transformers llm call · how do i balance context and history when creating prompts for llm's · what is embedding in llm · what is text embedding in llm · how can i teach a book to an llm · what is hybrid search in llm · what is reranking in llm

Also asked as: what is an llm · what is llm in ai · large language model meaning · llm meaning · what is a large language model · llm explained · what does llm stand for · what is llm model · what is generative ai

The architecture comes from the 2017 transformer paper [1]. The demonstration that scale alone unlocked few-shot learning came with GPT-3 [2]. The recipe for turning a text predictor into an assistant, supervised fine-tuning on instructions followed by preference training, is InstructGPT [5]. Everything shipping today descends from those three papers.

I build products on top of these models every week: RAG systems, fine-tuning platforms, agents, document extraction. This page is the mental model I give students and clients before any of that makes sense.

An LLM does not know things. It has a very good sense of what text comes next. Every product decision follows from taking that sentence seriously. Pranjul Rathour

How does an LLM work?

Text is split into tokens. Each token becomes a vector. Layers of attention let every token look at every earlier token and decide what matters. The final layer outputs a probability for every token in the vocabulary as the next one. Sampling picks one, appends it, and the loop repeats until a stop token. Training adjusts billions of weights so those probabilities match real text. Inference is the same forward pass, one token at a time, which is why output is streamed and why long outputs cost more.

Also asked as: how does an llm work · how llms work · how do large language models work · llm architecture · transformer explained · how does chatgpt work · how llm generates text · what is attention in llm · how are llms trained

Attention is the operation that made this work at scale: instead of compressing the past into a fixed state, every position can weigh every previous position directly [1]. It is also why cost grows with context length.

What is a token?

A token is the unit an LLM reads and writes: usually a word piece, sometimes a whole word or a single character, produced by a tokeniser trained with an algorithm such as byte-pair encoding [8]. In English, a token averages about three-quarters of a word; Hindi and code use more tokens per character. Prices, context limits and speed are all measured in tokens, so "how many words" questions are really token questions.

Also asked as: what is a token in llm · tokens vs words · how many tokens in a word · what is tokenization in llm · llm token limit · token count · how to count tokens · what is a context token

What are parameters in an LLM?

Parameters are the learned weights, the numbers adjusted during training. A 7B model has about seven billion. More parameters mean more capacity to store patterns, more memory to run, and more cost per token. Scaling laws showed that quality improves predictably with parameters and data [3], and the Chinchilla paper showed most models were undertrained for their size, which is why recent small models trained on far more tokens punch above their weight [4].

Also asked as: how to pass dynamic parameters (eg, index name, key) from semantic kernel to mcp server without llm decision-making · how many parameters for local llm

Also asked as: what are parameters in llm · llm parameters meaning · 7b vs 70b model · what does 7b mean · parameters vs tokens · llm size · small language model vs large language model · slm vs llm

What is a context window?

The context window is the maximum number of tokens the model can read and write in one request: the system prompt, the conversation, any documents you paste in, and the answer, all together. Windows have grown from a few thousand tokens to a million or more, but two things did not change: cost scales with what you put in, and models attend unevenly across long inputs, using the start and end better than the middle [7]. Retrieval exists to choose what deserves the window.

Also asked as: what is hallucination in the context of ai

Also asked as: what is context window in llm · context length vs context window · llm context window size · what happens when context window is full · long context llm · context window explained · how much text can an llm read

What is the difference between GPT, Claude, Gemini, Llama and Mistral?

They are families of LLMs from different organisations with different access models. GPT (OpenAI), Claude (Anthropic) and Gemini (Google) are closed: you call an API and never see the weights. Llama (Meta), Mistral, Gemma (Google) and Qwen (Alibaba) publish open weights you can download, run and fine-tune yourself. Quality differences between the top models are task-dependent and change with every release; the access model is the durable difference.

Also asked as: what is gemini google · what is google gemini ai · difference between hallucination and delusion · difference between hallucination and illusion · difference between hallucination and delusion and illusion · difference between hallucination and pseudohallucination · difference between hallucination and schizophrenia · difference between hallucination and delusion in hindi · difference between hallucination and imagination · difference between hallucination and delusion with example · difference between hallucination and delirium · difference between hallucination and dream · which google gemini api is free · difference between gemini api and google ai

Also asked as: gpt vs claude vs gemini · best llm · which llm is best · llama vs mistral · open source llm vs closed · gpt vs llama · best open source llm · claude vs chatgpt · gemini vs chatgpt · llm comparison

For rankings, use live leaderboards rather than any article: the Open LLM Leaderboard for open models on benchmarks [9] and LMArena for human preference across all models [10]. Then test the top three on your own task, because leaderboards are not your workload.

What is hallucination, and why do LLMs make things up?

Hallucination is fluent output that is false or unsupported: an invented citation, a wrong date, a function that does not exist. It happens because the model is predicting plausible text, not retrieving verified facts; when the training data is thin or the question is outside it, the most probable continuation is still produced. The survey by Ji et al. catalogues the causes and mitigations [6]. In practice you manage it with retrieval, so the model has evidence to quote, with constraints, so it can say "not in the documents", and with evaluation, so you measure how often it happens.

Also asked as: what is hallucination in ai · hallucination vs delusion · what does hallucination mean · what is hallucination in hindi · what is hallucination in psychology · what is hallucination means · what does hallucination mean in ai · what does hallucination mean in ai outputs · what does hallucination feel like · what does hallucination mean in demonology · what is considered a hallucination · what causes an hallucination · why is batman hallucination joker · why does hallucination happen in ai

Also asked as: what is hallucination in llm · why do llms hallucinate · ai hallucination examples · how to reduce hallucination · llm hallucination rate · can llms be trusted · why does chatgpt make things up · grounding llm

My RAG platform has a confidence gate for exactly this reason: when retrieved evidence is weak, it refuses to answer rather than produce a fluent guess. It is the feature clients noticed most, because a wrong answer with confidence is the most expensive kind.

What is the difference between an LLM and generative AI, machine learning and AI?

AI is the field. Machine learning is the approach of learning patterns from data instead of hand-coding rules. Deep learning is machine learning with large neural networks. Generative AI is deep learning models that produce content, text, images, audio, code. An LLM is a generative model for text. So an LLM is one kind of generative AI, which is one kind of deep learning, inside machine learning, inside AI. The nesting matters when someone asks whether you "know AI".

Also asked as: what is google ai · what is google ai called · what are embeddings in ai · best local ai · what is chunking in ai · can ai hallucination be regarded as a machine error · how to implement llm hallucination detection in production with guardrails ai 0.5 and langchain · what is the difference between llm and msc in law · how i upgrade xiaoai speaker local llm: sub-200ms ai · show hn: pokertools arena – local ai vs. ai poker llm benchmark table

Also asked as: llm vs generative ai · ai vs machine learning vs deep learning · what is generative ai vs ai · difference between llm and ai · generative ai vs llm · machine learning vs llm · is chatgpt an llm · is llm machine learning

How do I run an LLM locally?

Install Ollama, pull an open model sized for your machine, and run it from the terminal or its local API [11]. Under the hood it uses llama.cpp, which runs quantised GGUF models on CPUs and consumer GPUs [12]. A laptop with 8 GB of RAM runs 3 to 4B models comfortably; 16 GB runs 7 to 8B; a GPU with 24 GB runs 30B-class models quantised. For serving many users from a GPU server, vLLM is the standard [13].

Also asked as: show hn: run the popular llm-course tutorials on hyperai · how to setup llm locally · how to run an llm locally · can i run this llm · how to host llm locally · how to install llm locally · how can i run ollama in google colaboratory · how to locally access an ollama model remotely hosted on a google colab for a python script · what local llm can i run · how to run local llm on mac · how to run local llm on windows · how much to run local llm · can local llm run without internet · can i run local llm on macbook air m4

Also asked as: how to run llm locally · run llm on laptop · best local llm · ollama tutorial · llm on cpu · how much ram to run llm · run llm without gpu · local llm for coding · lm studio vs ollama · llama cpp

How much does an LLM cost to use?

API pricing is per million tokens, with separate rates for input and output, and output usually several times more expensive. Your cost per request is input tokens times the input rate plus output tokens times the output rate; your monthly cost is that times requests. A RAG system that stuffs ten passages into every prompt pays for those passages every time, which is why reranking to five and caching repeated prefixes matter. Prices change often; read the providers' pricing pages [14][15][16], and measure on real traffic before you quote a budget.

Also asked as: how much hallucination is normal · how much sleep deprivation till hallucination · how much vram for llm · how much ram for llm · how much google gemini api cost · can i use gemini api with google ai pro · can you use google gemini api for free · google veo 3.1 lite: build ai video apps at half the cost (2026 developer guide · how much ram local llm · how much vram local llm · how much does local llm cost · how much memory for local llm · how much ram for local llm reddit · how much ram for local llm mac

Also asked as: llm cost · how much does an llm cost · llm api pricing · token cost calculator · gpt api cost · claude api pricing · gemini api pricing · cost of running llm · how to reduce llm cost · llm cost optimization

Outages are a cost too. My systems fall back across providers with identical prompts when one API degrades, so the product stays up without paying for two calls on every request [17].

On the mic
On the mic

What can LLMs do, and what can they not do?

They are strong at transforming text: summarising, rewriting, translating, extracting, classifying, drafting and writing code from a clear specification. They are unreliable at arithmetic without tools, at facts outside their training, at knowing what they do not know, and at anything that needs a real-world action they cannot take. Products that work pair the model's strengths with tools and retrieval for its weaknesses.

Also asked as: google gemini api not working: quota errors, auth issues and rate limits explained

Also asked as: what can llms do · limitations of llms · what can chatgpt not do · llm use cases · llm applications · can llms reason · are llms intelligent · llm capabilities

What are LLM agents, RAG and fine-tuning, in one paragraph each?

RAG gives the model documents to read before answering, so facts come from your data with citations. Fine-tuning changes the model's weights on your examples, so its behaviour, style or skill changes. Agents put the model in a loop with tools, so it can act, observe and continue. Prompting sits under all three and is where every project should start. I have separate guides on each; the short version is knowledge, behaviour, action.

Also asked as: what are the best-practice architectural workflows for llm-based contract compliance agents

Also asked as: llm vs rag · rag vs fine tuning vs prompting · what is an llm agent · llm application architecture · how to build with llms · llm engineering

What are common LLM interview questions?

Explain what a token is and why it matters for cost. Explain attention in one minute. Define the context window and one failure mode of long contexts. Explain hallucination and two mitigations. Compare an open-weight model with an API model for a given product. Estimate the monthly cost of a feature. Say how you would evaluate an LLM output. Every one of these has an answer on this page.

Also asked as: python/genai/llm: best practices for asking follow‑up questions to clarify user intent in a chatbot · which hallucination is common in schizophrenia · which hallucination is most common in schizophrenia · which hallucination is common in alcohol withdrawal · which hallucination is common in alcohol · which type of hallucination is most common

Also asked as: llm interview questions · llm interview questions and answers · generative ai interview questions · transformer interview questions · nlp interview questions llm · ai interview questions for freshers

Where should I learn LLMs properly?

Read the three papers above with a notebook open, run an open model locally with Ollama, then build one product on top of it: a RAG system over documents you know is the best first project because it exercises tokens, context, prompting, hallucination and cost all at once. If your college wants a session that gets a room of students from "what is an LLM" to a working demo, email pranjulrathour41@gmail.com or use pranjulrathour.scult.in/invite.

Sources

  1. Vaswani et al., Attention Is All You Need (2017)arxiv.org
  2. Brown et al., Language Models are Few-Shot Learners (GPT-3, 2020)arxiv.org
  3. Kaplan et al., Scaling Laws for Neural Language Models (2020)arxiv.org
  4. Hoffmann et al., Training Compute-Optimal Large Language Models (Chinchilla, 2022)arxiv.org
  5. Ouyang et al., Training language models to follow instructions with human feedback (InstructGPT, 2022)arxiv.org
  6. Ji et al., Survey of Hallucination in Natural Language Generation (2022)arxiv.org
  7. Liu et al., Lost in the Middle: How Language Models Use Long Contexts (2023)arxiv.org
  8. Sennrich et al., Neural Machine Translation of Rare Words with Subword Units (BPE, 2016)arxiv.org
  9. Hugging Face, Open LLM Leaderboardhuggingface.co
  10. LMArena (Chatbot Arena) leaderboardlmarena.ai
  11. Ollama: run open models locallyollama.com
  12. llama.cppgithub.com
  13. vLLM documentationdocs.vllm.ai
  14. OpenAI API pricingopenai.com
  15. Anthropic API pricinganthropic.com
  16. Google Gemini API pricingai.google.dev
  17. Multi-provider LLM fallback: staying up when one API goes down, Pranjul Rathourpranjulrathour.scult.in
Pranjul Rathour
Pranjul Rathour
GenAI Engineer · Kanpur, Uttar Pradesh, India

GenAI engineer and AI product builder with 2+ years shipping production-grade AI systems: RAG pipelines, fine-tuned LLMs, hybrid retrieval and multi-modal apps across vision, speech and OCR, architected end to end from ingestion to deployment. Leads engineering for SCULT INDIA's 14-member team, founded the 500+ member TechVerse Enclave community and has mentored 200+ students. Three hackathon first prizes: Changethon 2025 (IIT Roorkee), Product Genesis at Vividhotsava 2025 (CSJMU Kanpur) and BYTEBATTLE (MeetKats).

Open to GenAI roles, hackathon judging, mentorship sessions and guest talks at colleges. Email pranjulrathour41@gmail.com.

Keep reading

More on RAG & retrieval

All RAG & retrieval guides →
Campus talks · hackathon judging · mentoring

Want this as a live session at your college?

Open to GenAI roles, hackathon judging, mentorship sessions and guest talks at colleges.

Pranjul Rathour, GenAI engineer in Kanpur