Key takeaways
- Loading a model with transformers and calling generate() works for one request at a time; it falls over under real concurrent traffic.
- vLLM's PagedAttention and continuous batching are what let dozens of concurrent requests share a GPU efficiently.
- Ollama and llama.cpp optimise for easy local use on a laptop, not maximum throughput on a server; they solve a different problem well.
- SGLang adds structured generation and faster serving for certain workloads, particularly ones with repeated prompt structure.
- Choose by your actual constraint: concurrent users and throughput point to vLLM or SGLang; a laptop or a single user points to Ollama or llama.cpp.
What is vLLM?
vLLM is an open-source inference server for large language models built around PagedAttention, a memory-management technique that lets many concurrent requests share a GPU's memory efficiently by managing the attention key-value cache the way an operating system manages virtual memory, in pages, rather than reserving one large contiguous block per request [1][2]. The practical result is much higher throughput on the same hardware compared with naively calling a model's generate() method for each request, which is why vLLM is the default choice for serving open-weight models at real request volume.
Also asked as: what is vllm · vllm explained · vllm meaning · how does vllm work · vllm vs transformers · what is pagedattention · vllm architecture · why use vllm
Every production system I run that serves an open model sits on some form of batched serving; the difference between "it works in a notebook" and "it serves five hundred requests a minute without falling over" is almost entirely this layer.
A model that answers one request beautifully and a service that answers a thousand requests a minute are different engineering problems. vLLM, SGLang and friends are the answer to the second one. Pranjul Rathour
Why can't I just call model.generate() for production traffic?
Because it processes one request, or a small fixed batch, at a time, and each request occupies GPU memory for its full duration even while it is producing tokens slowly one at a time. Real traffic arrives as many concurrent requests of different lengths, and naive batching either wastes memory reserving worst-case space for every request or wastes GPU time waiting for a batch to fill before starting. Continuous batching solves the second problem, starting and finishing requests within a batch as they arrive and complete rather than waiting for a fixed batch to fill [8]. PagedAttention solves the first, sharing memory pages between requests instead of reserving a contiguous block per request [1].
Also asked as: why is generate slow for production · naive llm serving problems · continuous batching explained · llm serving bottleneck · why does llm inference need special serving · transformers generate vs vllm
What is Ollama, and how is it different from vLLM?
Ollama is a tool for running open-weight models easily on a laptop or a single machine, wrapping llama.cpp with a simple CLI, a model library, and an API that feels like calling a hosted service [3]. It optimises for ease of local use, quantized models that fit on consumer hardware, and a good single-user experience, not maximum concurrent throughput. Running Ollama to serve many simultaneous users at scale is possible but is not what it is built or tuned for; vLLM is built specifically for that case.
Also asked as: ollama vs vllm · is ollama free · what is ollama · ollama for production · ollama vs llama.cpp · can ollama serve multiple users · ollama scalability
What is llama.cpp, and where does it fit?
llama.cpp is a C++ inference engine for running LLMs efficiently on CPUs and consumer GPUs, using quantized model formats such as GGUF to shrink memory use [4]. It underlies Ollama and many other local-inference tools. Its strength is portability and running well on hardware without a data-centre GPU, phones, laptops, Raspberry Pi-class devices; it is not designed around maximising throughput for many concurrent server requests the way vLLM is.
Also asked as: what is llama.cpp · llama.cpp vs vllm · llama.cpp explained · gguf and llama.cpp · how does llama.cpp work · llama.cpp on cpu

What is SGLang, and how does it compare to vLLM?
SGLang is an inference framework that combines fast serving, similar in spirit to vLLM's continuous batching and paged memory, with a structured-generation front end that lets you express constrained outputs, branching, and reused prompt prefixes efficiently [5][6]. Its RadixAttention technique caches shared prompt prefixes across requests, which is particularly effective for workloads with repeated system prompts or few-shot examples, agents calling the same tools repeatedly. Benchmarks between vLLM and SGLang shift with every release of both; the honest answer is to benchmark your own workload rather than trust either project's numbers alone.
Also asked as: what is sglang · sglang vs vllm · sglang explained · radixattention · structured generation llm serving · sglang benchmark
How do I decide what I actually need?
Count your expected concurrent requests per second and your latency budget, then pick backward from that. A weekend project or an internal tool with a handful of users rarely needs more than Ollama on a single GPU or even a good CPU. A product with real concurrent traffic needs vLLM or SGLang on dedicated GPU serving from day one, because retrofitting proper batching after launch under load is a much worse day than provisioning for it up front.
Also asked as: how to choose llm serving framework · how many requests can vllm handle · llm serving capacity planning · when do i need vllm · llm inference sizing · production llm serving checklist
What does this cost, and how does quantization fit in?
Serving throughput and quantization are separate levers that stack: a quantized model uses less memory and can run on smaller GPUs, while a serving engine like vLLM uses that freed memory to run more concurrent requests rather than only fitting a bigger model [9]. The cheapest production setup is usually a quantized model served through vLLM or SGLang on the smallest GPU that hits your latency target, re-benchmarked whenever a new quantization method or a smaller strong model appears, since the frontier moves every few months.
Also asked as: vllm quantization · cheapest way to serve llm · llm serving cost · gpu sizing for llm serving · vllm gpu requirements · how much gpu memory does vllm need
LLM serving interview questions
Explain what problem continuous batching solves. Describe PagedAttention in your own words. Compare vLLM, SGLang, Ollama and llama.cpp and say when you would choose each. Explain how quantization and serving throughput interact. Describe how you would capacity-plan a serving setup for a given requests-per-second target. The strongest answer includes a real throughput number you measured, not a number you read.
Also asked as: llm serving interview questions · vllm interview · ai infrastructure interview questions · llm inference optimization interview
Where should I start?
Stand up vLLM on a rented GPU with a small open model and hit it with a simple load-testing script, then repeat with Ollama on the same hardware and compare throughput yourself. The gap will be obvious within the first ten minutes. For a hands-on session on serving LLMs in production, from a notebook to real concurrent traffic, email pranjulrathour41@gmail.com or use pranjulrathour.scult.in/invite.
Sources
- Kwon et al., Efficient Memory Management for Large Language Model Serving with PagedAttention (vLLM, 2023)arxiv.org
- vLLM documentationdocs.vllm.ai
- Ollama documentationgithub.com
- ggml-org, llama.cppgithub.com
- SGLang documentationdocs.sglang.ai
- Zheng et al., SGLang: Efficient Execution of Structured Language Model Programs (2024)arxiv.org
- Hugging Face, Text Generation Inference (TGI)huggingface.co
- NVIDIA, continuous batching explaineddeveloper.nvidia.com
- What is quantization in LLMs, Pranjul Rathourpranjulrathour.github.io
- How much does an LLM API cost, Pranjul Rathourpranjulrathour.github.io




