Key takeaways
- Quantization replaces 16-bit weights with 8-bit or 4-bit approximations, cutting memory by two to four times with a small, measurable accuracy cost.
- GGUF is the format for llama.cpp on CPUs and consumer GPUs; AWQ and GPTQ are GPU-oriented methods for serving; NF4 is what QLoRA uses during training.
- 8-bit is nearly lossless; 4-bit is usually fine for chat and summarisation and worse for maths, code and long reasoning. Measure on your task.
- Quantize after fine-tuning and merging, never before, unless you are doing QLoRA where the base is quantized on purpose.
- Memory in gigabytes is roughly parameters in billions times bytes per weight, plus overhead. A 7B model is about 14 GB at 16-bit, about 4 to 5 GB at 4-bit.
What is quantization in LLMs?
Quantization is storing a model's weights, and sometimes its activations, in fewer bits than they were trained in: 8-bit or 4-bit integers instead of 16-bit floats. Each weight is mapped to a small grid of values with a scale factor per block, so the model takes a half or a quarter of the memory and often runs faster, at the cost of a small approximation error. It is why a 7-billion-parameter model that needs about 14 GB in 16-bit runs on a laptop with 8 GB in 4-bit.
Also asked as: what is quantization in llm · quantization in machine learning · llm quantization explained · what is quantization in deep learning · 4 bit quantization llm · 8 bit quantization · what does quantized model mean · quantized llm meaning · model quantization
Quantization is the reason open models became usable outside data centres. LLM.int8() showed 8-bit inference at scale with almost no loss [1]; the k-bit scaling-law paper argued that for a fixed memory budget, 4-bit weights on a bigger model usually beat 16-bit weights on a smaller one [5]. QLoRA then made 4-bit the basis for training, not just inference [2].
I quantize every model that leaves FineTune Studio, because a fine-tuned adapter is only useful once it is running somewhere, and "somewhere" is usually a machine smaller than the one that trained it [12].
Quantization is the difference between a model that impressed you in a notebook and a model your users can actually run. Pranjul Rathour
How does quantization work?
Weights are grouped into blocks, typically 32 to 128 values. For each block, the largest magnitude sets a scale, and every weight is rounded to the nearest point on a small integer grid, 256 points for 8-bit, 16 for 4-bit. At inference, the integers are multiplied back by the scale, exactly or approximately, before the matrix multiplication. Smarter methods choose the grid or the rounding to minimise error on real activations rather than on the weights alone.
Also asked as: how does quantization work · quantization algorithm · block quantization · symmetric vs asymmetric quantization · quantization scale and zero point · weight quantization vs activation quantization · how quantized models run
What is NF4?
NF4, NormalFloat4, is the 4-bit data type QLoRA introduced: instead of an evenly spaced grid, its 16 levels are placed where normally distributed weights actually cluster, so the same 4 bits capture more of the distribution [2]. It is a training-time quantization for the frozen base model, held in memory while adapters train; it is not a serving format. When people fine-tune "in 4-bit", this is usually what they mean.
Also asked as: what is nf4 · nf4 vs int4 · normalfloat4 · qlora quantization type · bnb 4bit quant type · double quantization qlora
What are GGUF, AWQ and GPTQ?
GGUF is a file format used by llama.cpp that packages a quantized model with its metadata for fast loading on CPUs, Apple Silicon and consumer GPUs; its quantization schemes are named like Q4_K_M and Q8_0 [6]. GPTQ is a post-training method that quantizes weights layer by layer using calibration data to minimise output error, aimed at GPU inference [3]. AWQ observes that a small fraction of weights matter far more, protects them by scaling, and quantizes the rest, also for GPUs [4]. Same goal, different trade-offs and runtimes.
Also asked as: gguf vs awq vs gptq · what is gguf · gguf meaning · awq vs gptq · gptq vs gguf · what is awq quantization · q4_k_m meaning · gguf quantization types · which quantization format to use · exl2 vs gguf
The rule of thumb I give: GGUF if it runs on a laptop or a phone, AWQ or GPTQ if it runs on a GPU server behind vLLM [10], NF4 through bitsandbytes if you are training [7]. The Transformers quantization guide covers loading each [8].
How much quality do you lose with quantization?
Little at 8-bit, usually within noise on most benchmarks. Modest at 4-bit for chat, summarisation and extraction, and more noticeable on arithmetic, code generation and long multi-step reasoning, where small errors compound. Below 4 bits the drop steepens quickly. Larger models tolerate quantization better than small ones, so a 4-bit 13B usually beats a 16-bit 7B in the same memory [5]. Measure on your own task before trusting any of these generalisations.
Also asked as: quantization accuracy loss · does quantization reduce quality · 4 bit vs 8 bit quality · quantization perplexity · is 4 bit quantization good · quantization impact on performance · q4 vs q8 · how much accuracy lost in quantization
How is quantization related to QLoRA and fine-tuning?
QLoRA quantizes the frozen base model to NF4 so it fits in memory while small adapters train in higher precision [2]. That is quantization for training. After training you merge the adapter into the 16-bit base weights and then quantize the merged model for serving with GGUF, AWQ or GPTQ. Do not quantize before merging, and do not merge into a 4-bit base for serving; reload the base in 16-bit, merge, then quantize.
Also asked as: qlora quantization · quantize after fine tuning · merge lora then quantize · how to convert fine tuned model to gguf · quantize fine tuned model · lora merge quantization · can you fine tune a quantized model
My own notes on this pipeline, with the commands, are on the portfolio [11].

How much memory does a quantized model need?
Roughly parameters in billions times bytes per weight, plus a few gigabytes for activations and the KV cache that grows with context length and concurrent users. Two bytes per weight at 16-bit, one at 8-bit, about half a byte at 4-bit. So 7B is about 14, 7 and 4 to 5 GB; 13B about 26, 13 and 8 GB; 70B about 140, 70 and 40 GB. Long contexts add memory fast, which is why a model that "fits" can still run out at 32,000 tokens.
Also asked as: how much vram for 7b model · how much ram to run llm · llm memory requirements · 13b model vram · 70b model requirements · vram calculator llm · how much memory does a quantized model need · kv cache memory
How do I quantize a model myself?
For local use: convert the merged model to GGUF with llama.cpp's conversion script and quantize to a variant such as Q4_K_M or Q8_0, then run it with llama.cpp or Ollama [6][9]. For GPU serving: quantize with AWQ or GPTQ using a few hundred calibration samples that resemble your real inputs, then serve with vLLM [10]. For loading in Transformers without a conversion step: bitsandbytes 8-bit or 4-bit at load time [7]. Always keep the 16-bit merged model as the source of truth.
Also asked as: how to quantize a model · how to quantize llm · convert to gguf · llama.cpp quantize · how to make gguf file · awq quantization tutorial · gptq quantization tutorial · bitsandbytes 4bit loading · quantize model for ollama
Which quantization should I choose?
Decide by where the model runs and what it does. Laptop or phone: GGUF Q4_K_M as the default, Q8_0 if memory allows. GPU server: AWQ or GPTQ 4-bit behind vLLM, 8-bit if quality on hard tasks matters. Training: NF4 via bitsandbytes. Maths, code or long reasoning as the main task: prefer 8-bit or a bigger 4-bit model. Then confirm on your eval set. The wrong answer is to skip the measurement.
Also asked as: best quantization for llm · which quantization to use · q4_k_m vs q5_k_m vs q8_0 · 4 bit or 8 bit · best gguf quant · quantization for coding models · quantization for local llm
Quantization interview questions
Explain what quantization does and why it saves memory. Compare 8-bit and 4-bit loss. Explain the difference between GGUF, AWQ and GPTQ and where each runs. Explain NF4 and its role in QLoRA. Describe the adapter-to-serving pipeline. Estimate the memory of a 13B model at 4-bit. Each is answered above; the strongest evidence is a model you quantized and measured.
Also asked as: quantization interview questions · llm inference optimization interview · model compression interview questions · qlora interview questions
Where should I start?
Pull a 7B model in GGUF through Ollama at Q4_K_M and at Q8_0, ask each the same twenty questions including some arithmetic and code, and note where they differ. Then quantize a model you fine-tuned yourself and run the same comparison against the 16-bit merged version. For a hands-on session on fine-tuning and deploying models on modest hardware, email pranjulrathour41@gmail.com or use pranjulrathour.scult.in/invite.
Sources
- Dettmers et al., LLM.int8(): 8-bit Matrix Multiplication for Transformers at Scale (2022)arxiv.org
- Dettmers et al., QLoRA: Efficient Finetuning of Quantized LLMs (2023)arxiv.org
- Frantar et al., GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers (2022)arxiv.org
- Lin et al., AWQ: Activation-aware Weight Quantization for LLM Compression and Acceleration (2023)arxiv.org
- Dettmers & Zettlemoyer, The case for 4-bit precision: k-bit Inference Scaling Laws (2022)arxiv.org
- llama.cpp and the GGUF formatgithub.com
- bitsandbytes documentationhuggingface.co
- Hugging Face Transformers quantization guidehuggingface.co
- Ollama: run open models locallyollama.com
- vLLM quantization documentationdocs.vllm.ai
- Quantization after fine-tuning: GGUF, AWQ, GPTQ, Pranjul Rathourpranjulrathour.scult.in
- FineTune Studio source code, Pranjul Rathourgithub.com



