Key takeaways
- LoRA trains small adapter matrices on a frozen base; QLoRA does the same on a 4-bit quantised base; full fine-tuning updates every weight.
- Memory is the deciding factor: QLoRA puts a 7B model on a free 16 GB GPU and a 1 to 3B model on a laptop; full fine-tuning of 7B needs a cluster.
- Quality: LoRA is close to full fine-tuning for most behaviour tasks; QLoRA is close to LoRA. Measure on your task, not on a leaderboard.
- Rank, alpha, learning rate, epochs and target modules are the settings that matter. Apply adapters to all linear layers unless memory forbids.
- Choose QLoRA by default on consumer hardware, LoRA when you have the memory and want speed, full fine-tuning only with a cluster and a reason.
LoRA vs QLoRA vs full fine-tuning: what is the difference?
Full fine-tuning updates every weight in the model, which needs memory for the weights, their gradients and optimiser states, several times the model's size. LoRA freezes the weights and trains two small low-rank matrices per targeted layer whose product approximates the update, cutting trainable parameters to well under one percent [1]. QLoRA keeps the LoRA adapters but stores the frozen base in 4-bit NormalFloat, so the base itself takes a quarter of the memory [2]. Same adapters, same training loop, radically different memory.
Also asked as: lora vs qlora · qlora vs lora · lora vs full fine tuning · qlora vs full fine tuning · difference between lora and qlora · lora qlora full fine tuning comparison · what is lora and qlora · peft vs full fine tuning · lora or qlora which is better · lora vs qlora memory
FineTune Studio, the platform I built in 2026, runs QLoRA from the browser precisely because QLoRA is what makes fine-tuning fit on hardware students and small teams actually have [13]. My own run of a 1.7-billion-parameter model in about 3.2 GB of VRAM is the concrete number behind this page, and the longer comparison notes are on the portfolio [11][12].
Full fine-tuning is a decision about hardware. LoRA is a decision about speed. QLoRA is the decision that lets you start today. Pranjul Rathour, from building FineTune Studio
How much memory does each method need?
Rough orders of magnitude for a fine-tuning run, before activations: full fine-tuning needs about 16 bytes per parameter with a standard optimiser, so a 7B model needs over 100 GB. LoRA holds the base in 16-bit, 2 bytes per parameter, plus small adapters and their optimiser state, so 7B is about 16 to 20 GB. QLoRA holds the base in 4-bit, about half a byte per parameter, so 7B is about 6 to 10 GB with activations at moderate sequence lengths. These move with sequence length, batch size and gradient checkpointing.
Also asked as: lora memory requirements · qlora vram · how much vram for lora · how much vram for qlora · vram for full fine tuning · fine tune 7b model vram · fine tune 13b model requirements · fine tune 70b qlora · qlora 7b 16gb · vram calculator fine tuning
These are planning numbers, not promises. The QLoRA paper's headline was a 65B model fine-tuned on a single 48 GB GPU [2]; my VRAM walkthrough for consumer cards has the measured figures for small models [12].
How do speed and quality compare?
Speed: LoRA is fastest per step, because the base runs in 16-bit; QLoRA is slower per step, because 4-bit weights are de-quantised on the fly, though it often wins overall by allowing larger batches on the same GPU; full fine-tuning is slowest and most expensive. Quality: LoRA matches full fine-tuning on most behaviour tasks and lags on tasks needing large capacity changes; Biderman et al. found LoRA learns less but also forgets less [4]. QLoRA matches LoRA closely when adapters target all linear layers [2].
Also asked as: lora vs qlora speed · qlora slower than lora · lora vs full fine tuning quality · does qlora reduce quality · qlora accuracy loss · lora performance vs full fine tuning · lora forgets less · fine tuning quality comparison
Lialin et al.'s survey places these among the wider family of parameter-efficient methods [3]; DoRA is a recent refinement that decomposes weights into magnitude and direction and often edges out LoRA at the same rank [5].
Which should I choose?
QLoRA if you are on a free Colab GPU, a laptop, or any card under 24 GB, or if you want the largest model that fits. LoRA if you have the memory and want faster iteration, for example a 24 GB or larger GPU with a 7B model. Full fine-tuning only with a multi-GPU setup, a large dataset and evidence that adapters fall short on your task. For almost every student and startup project, the answer is QLoRA.
Also asked as: should i use lora or qlora · when to use qlora · when to use lora · when to use full fine tuning · lora or qlora for 7b · best fine tuning method for small gpu · fine tuning method for beginners · qlora for consumer gpu
What settings actually matter for LoRA and QLoRA?
Five. Rank, the adapter's inner dimension: 8 to 32 for most tasks, 64 for demanding ones. Alpha, the scaling: equal to rank or twice it; the ratio sets the adapter's effective learning rate. Target modules: all linear layers, attention and feed-forward, unless memory forbids; the QLoRA paper found this mattered more than rank [2]. Learning rate: around 2e-4 for adapters, ten times higher than full fine-tuning. Epochs: one to three on small datasets; more is memorising. Dropout of 0.05 is a fine default.
Also asked as: lora rank · lora alpha · lora rank vs alpha · lora target modules · best lora rank · lora hyperparameters · qlora hyperparameters · lora learning rate · lora dropout · r and alpha in lora · all linear target modules
How do I fine-tune on a free Colab GPU?
Use QLoRA on a 1 to 7B instruct model with sequence length trimmed to your data, gradient checkpointing, a small per-device batch with accumulation, and either PEFT plus TRL or Unsloth, which cuts memory and time further [6][7][9]. Save adapter checkpoints to Drive or the Hub every few hundred steps because sessions end. A 7B QLoRA run at short sequence lengths fits a free T4; a 1 to 3B run leaves headroom [10].
Also asked as: fine tune llm on colab free · qlora colab · fine tune 7b on t4 · fine tune llm free gpu · unsloth colab · fine tune llama on colab · fine tune mistral free · how to fine tune on google colab · colab out of memory fine tuning
from transformers import AutoModelForCausalLM, AutoTokenizer, BitsAndBytesConfig
from peft import LoraConfig
from trl import SFTTrainer, SFTConfig
bnb = BitsAndBytesConfig(load_in_4bit=True, bnb_4bit_quant_type="nf4",
bnb_4bit_use_double_quant=True, bnb_4bit_compute_dtype="bfloat16")
model = AutoModelForCausalLM.from_pretrained(BASE, quantization_config=bnb, device_map="auto")
tok = AutoTokenizer.from_pretrained(BASE)
lora = LoraConfig(r=16, lora_alpha=32, lora_dropout=0.05, target_modules="all-linear", task_type="CAUSAL_LM")
cfg = SFTConfig(output_dir="out", num_train_epochs=2, learning_rate=2e-4, per_device_train_batch_size=2,
gradient_accumulation_steps=8, gradient_checkpointing=True, max_length=1024,
eval_strategy="steps", eval_steps=50, save_steps=200, logging_steps=10, bf16=True)
SFTTrainer(model=model, args=cfg, train_dataset=train, eval_dataset=val, processing_class=tok, peft_config=lora).train()
Argument names change between library versions; check the current PEFT, TRL and bitsandbytes docs before copying [6][7][8]. If memory runs out, in order: shorter max length, batch size 1 with more accumulation, a smaller model.

What do I do with the adapter after training?
Either keep it separate and load it on the base at inference, which lets one base serve many adapters, or merge it into the base weights to produce a single model. For QLoRA, merge into a freshly loaded 16-bit base, not the 4-bit one, then quantise the merged model to GGUF, AWQ or GPTQ for serving. Always publish the chat template and the training configuration with the adapter, or nobody, including you next month, can use it correctly.
Also asked as: merge lora adapter · lora adapter vs merged model · how to use lora adapter · serve lora adapters · multi lora serving · merge qlora adapter · lora adapter hugging face · save lora adapter
Common LoRA and QLoRA mistakes
Targeting only attention layers when memory allowed all linear layers. Rank 256 because bigger felt safer. A learning rate copied from full fine-tuning, ten times too low. Ten epochs on five hundred examples. Merging into the 4-bit base. Forgetting the chat template at serving time. Never comparing with the base model. Each is fixed by the checklist above and by the habit FineTune Studio enforces: evaluate base versus tuned side by side before anything ships [13].
Also asked as: lora mistakes · qlora not working · lora fine tuning not improving · lora overfitting · qlora out of memory · lora training tips · common fine tuning errors · lora troubleshooting
LoRA and QLoRA interview questions
Explain what LoRA changes and why it saves memory. Explain what QLoRA adds and the trade-off in speed. Compare quality and forgetting between the three methods. Name the settings that matter and defend a starting configuration. Describe the memory of a 7B run under each method. Explain how to merge and serve a QLoRA adapter correctly. An adapter you trained, with a base-versus-tuned comparison, is the strongest evidence.
Also asked as: lora interview questions · qlora interview questions · peft interview questions · fine tuning interview questions lora · explain lora in interview
Where should I start?
Fine-tune a 1 to 2B instruct model with QLoRA on Colab tonight, on two hundred examples you wrote, and compare it with the base on twenty held-out prompts. Then repeat with LoRA if your GPU allows and compare speed. The difference between the methods is best learned by watching the memory graph. For a hands-on fine-tuning workshop at your college, email pranjulrathour41@gmail.com or use pranjulrathour.scult.in/invite.
Sources
- Hu et al., LoRA: Low-Rank Adaptation of Large Language Models (2021)arxiv.org
- Dettmers et al., QLoRA: Efficient Finetuning of Quantized LLMs (2023)arxiv.org
- Lialin et al., Scaling Down to Scale Up: A Guide to Parameter-Efficient Fine-Tuning (2023)arxiv.org
- Biderman et al., LoRA Learns Less and Forgets Less (2024)arxiv.org
- Liu et al., DoRA: Weight-Decomposed Low-Rank Adaptation (2024)arxiv.org
- Hugging Face PEFT documentationhuggingface.co
- Hugging Face TRL documentationhuggingface.co
- bitsandbytes documentationhuggingface.co
- Unslothgithub.com
- Google Colabcolab.research.google.com
- LoRA vs QLoRA vs full fine-tuning, Pranjul Rathourpranjulrathour.scult.in
- Fine-tuning on a consumer GPU: a VRAM budget walkthrough, Pranjul Rathourpranjulrathour.scult.in
- FineTune Studio source code, Pranjul Rathourgithub.com




