How to fine-tune an LLM: LoRA, QLoRA, datasets and evaluation, from building a fine-tuning platform

What fine-tuning an LLM means, when it beats RAG or prompting, how LoRA and QLoRA make it fit on a free GPU, how to prepare a dataset, read the loss curve and evaluate honestly.

Presenting to a room
Presenting to a room

Key takeaways

  • Fine-tuning changes a model's weights to change its behaviour, style, format or skill. It does not reliably add facts; use RAG for facts.
  • LoRA trains small adapter matrices instead of all weights; QLoRA adds 4-bit quantisation of the frozen base, which is how a 7B model fits a free 16 GB GPU.
  • The dataset is the product. A few hundred to a few thousand clean, consistent examples in a chat template beat a hundred thousand scraped ones.
  • Read the loss curve for the gap between training and validation loss, stop at the plateau, and evaluate base versus fine-tuned on held-out prompts side by side.
  • Rank, alpha, learning rate and epochs are the hyper-parameters that matter; everything else is defaults until you have a reason.

What is fine-tuning an LLM?

Fine-tuning is continuing to train a pre-trained language model on your own examples so that its weights shift toward your task: your tone, your output format, your domain's vocabulary, or a narrow skill the base model does badly. You start from a model that already understands language and nudge it, which is why a few thousand good examples can change behaviour that no prompt could.

Also asked as: what does fine tuning llm mean · what is fine tuning in ai · what is fine tuning in machine learning · fine tuning meaning · what is llm fine tuning · what is supervised fine tuning

The base model learned from trillions of tokens of general text. Fine-tuning shows it a comparatively tiny, curated set of input and output pairs and adjusts the weights to make those outputs more likely. The result is a model that behaves differently by default, without a long prompt reminding it every time.

I built FineTune Studio in 2026 to turn this into a product workflow: upload a dataset, pick a base model, set LoRA or QLoRA hyper-parameters, launch a run, and watch the loss stream to the browser in real time [15]. Most of what follows is what that platform taught me about which decisions matter.

Fine-tuning is not how you teach a model facts. It is how you teach a model habits. If you remember one line from this page, make it that one. Pranjul Rathour, from building FineTune Studio

How does fine-tuning an LLM work?

Training runs the model on each example, compares its predicted next tokens with the target answer, computes a loss, and adjusts weights by gradient descent to lower that loss. Full fine-tuning updates every weight, which needs enormous memory. Parameter-efficient methods freeze the base model and train a small set of added weights instead, which is what almost everyone outside a large lab does today.

Also asked as: how does fine tuning llm work · how fine tuning works in llm · fine tuning process llm · how to fine tuning llm · llm fine tuning explained

Lialin et al. survey the parameter-efficient family, adapters, prompt tuning, LoRA and their relatives, and the trade-offs between them [3]. In practice the field has converged: LoRA when memory allows, QLoRA when it does not.

When should I fine-tune instead of using RAG or prompting?

Fine-tune when the problem is behaviour: consistent style, strict structured output, a domain's phrasing, a classification or extraction skill, or when you need a small model to do one thing well and cheaply. Use RAG when the problem is knowledge that changes or must be cited. Use prompting first in every case, because a prompt is free to change and a fine-tune is not.

Also asked as: rag vs fine tuning · fine tuning vs prompt engineering · when to fine tune llm · fine tuning vs rag vs prompt engineering · should i fine tune an llm · fine tuning vs few shot

The order I insist on with clients: prompt, then retrieval, then few-shot examples in the prompt, then fine-tuning. Each step is more expensive and less reversible than the last. Teams that fine-tune first usually end up fine-tuning again a month later when the requirements move.

Does fine-tuning add knowledge to a model?

Only weakly and unreliably. Fine-tuning on facts can make a model repeat those facts, but it cannot cite them, it forgets them as the world changes, and it tends to blend them with what it already believed. If the answer must be correct and traceable, retrieve it. LIMA showed that a thousand carefully chosen examples were enough to teach strong instruction-following, and that the knowledge came almost entirely from pre-training [6]. That is the clearest evidence for "fine-tune for behaviour, retrieve for facts".

Also asked as: can fine tuning add knowledge · does fine tuning teach new information · fine tuning for knowledge injection · fine tuning vs rag for knowledge

What is LoRA?

LoRA, low-rank adaptation, freezes the original model weights and trains two small matrices per targeted layer whose product approximates the weight update. Instead of updating billions of parameters you train a few million, the adapter, which is a few megabytes on disk and can be swapped or merged. Hu et al. introduced it in 2021 and showed quality comparable to full fine-tuning at a fraction of the memory [1].

Also asked as: what is lora in ai · what is lora fine tuning · lora meaning in machine learning · how does lora work · lora explained · what is lora adapter

The word LoRA also names a long-range radio protocol, which is why half the search results for "what is LoRA" are about antennas. In machine learning it is always low-rank adaptation.

Two settings define an adapter:

  • Rank (r). The inner dimension of the two matrices. Higher rank means more capacity and more memory. Typical values run from 8 to 64; I start at 16.
  • Alpha. A scaling factor applied to the adapter's output. A common rule is alpha equal to rank or twice the rank; what matters is the ratio, since it changes the effective learning rate of the adapter.

What are the target modules?

Target modules are the layers inside the transformer that get an adapter. Attention projections (query, key, value, output) are the classic choice; adding the feed-forward projections usually improves results at some memory cost. The QLoRA paper found applying adapters to all linear layers mattered more than rank for matching full fine-tuning [2].

Also asked as: lora target modules · which layers to apply lora · lora rank and alpha · lora hyperparameters

What is QLoRA, and what is the difference between LoRA and QLoRA?

QLoRA is LoRA on top of a base model quantised to 4 bits. The frozen weights are stored in a compact NormalFloat4 format and de-quantised on the fly during the forward pass, while the small adapters train in higher precision. Dettmers et al. showed a 65B model fine-tuning on a single 48 GB GPU with quality matching 16-bit full fine-tuning [2]. The difference from LoRA is memory: same adapters, much smaller base.

Also asked as: what is qlora · what is lora and qlora · lora vs qlora · difference between lora and qlora · qlora explained · qlora vs full fine tuning

The memory figures are orders of magnitude, not promises; they move with sequence length, batch size and gradient checkpointing. The reason QLoRA matters to students is simple: it is the technique that puts a 7B model on a free Colab T4 or a 16 GB laptop GPU. FineTune Studio's whole premise is that fine-tuning is a browser workflow, not a cluster job, and QLoRA is why that premise holds [15].

What is quantisation, and what are GGUF, AWQ and GPTQ?

Quantisation stores weights in fewer bits, 8 or 4 instead of 16, which shrinks memory and speeds inference at a small accuracy cost. GGUF is the file format used by llama.cpp for running quantised models on CPUs and consumer GPUs [13]; AWQ and GPTQ are post-training quantisation methods aimed at GPU inference. QLoRA's NormalFloat4 is a quantisation for training. Quantise after fine-tuning and merging, not before, unless you are doing QLoRA.

Also asked as: what is quantization in llm · gguf vs awq vs gptq · how to quantize a fine tuned model · 4 bit quantization llm · what is gguf

How do I prepare a dataset for fine-tuning?

Write the behaviour you want as a spec, then collect or write examples that follow it exactly, in the chat format the base model expects: a system message, a user turn and the ideal assistant answer. Deduplicate, validate every row, hold out ten to twenty percent for evaluation, and read a random sample by eye before training. Consistency beats size; a thousand examples that all follow the spec outperform fifty thousand that mostly do.

Also asked as: how to prepare dataset for fine tuning · fine tuning dataset format · best fine tuning datasets · how much data to fine tune llm · alpaca format vs sharegpt format · chat template fine tuning · how to create dataset for llm

Alpaca popularised the instruction, input, output format with 52,000 generated examples [8]; ShareGPT-style multi-turn conversations followed. Both are just JSON shapes. What the model actually sees is the rendered chat template, so use the tokenizer's template function rather than hand-writing the special tokens [11]. Mismatched templates are the most common reason a fine-tuned chat model produces garbage.

How much data do I need to fine-tune an LLM?

For style, tone and format, a few hundred high-quality examples produce a visible change. For a classification or extraction skill, a few thousand. For a new domain's language, tens of thousands. Below a hundred, prompting with examples usually does as well. LIMA's thousand examples are the reference point for how little alignment data can be if it is clean [6].

Also asked as: how much data to fine tune llm · minimum dataset size for fine tuning · how many examples for fine tuning gpt · fine tuning with small dataset

Can I use synthetic data for fine-tuning?

Yes, and most teams do, by having a stronger model generate examples from a spec or from seed documents. It works when you filter hard: validate every synthetic example against the spec, deduplicate, and mix in real examples so the model does not learn the generator's tics. Alpaca itself was trained on synthetic data generated from 175 human-written seeds [8].

Also asked as: synthetic data for fine tuning · how to generate training data with llm · is synthetic data good for fine tuning

Which base model should I fine-tune?

Pick the smallest open-weight model that already does the task acceptably with a prompt, in a size that fits your GPU with QLoRA, with a licence you can ship under. Start from the instruct variant when your task is conversational and from the base variant when you are teaching a completion format from scratch. Small models, 1 to 4 billion parameters, fine-tune faster, serve cheaper and are often enough.

Also asked as: best llm to fine tune · best open source model for fine tuning · base model vs instruct model fine tuning · which model to fine tune for chatbot · small language models fine tuning · can gemma be fine tuned · fine tune llama vs mistral

In my own write-up of a run on FineTune Studio, a 1.7-billion-parameter model fine-tuned with QLoRA in about 3.2 GB of VRAM [16]. That is laptop territory. The trade-off you accept is ceiling: a small model learns a narrow behaviour well and generalises less. For a support-ticket classifier that is exactly right. For an open-ended assistant it is not.

Can I fine-tune GPT or other closed models?

Yes, through the vendor's API. OpenAI offers fine-tuning of selected models with an upload-and-train workflow, priced per training and inference token [14]. You get no weights, no adapter file and no local inference; you get a model identifier. It is the right choice when you already depend on that vendor and do not want to run inference yourself, and the wrong one when you need control, portability or offline use.

Also asked as: how to fine tune gpt · can i fine tune gpt 4 · how to fine tune gpt 4o · fine tune chatgpt on my data · openai fine tuning cost · how to fine tune claude

How do I fine-tune an LLM on a free GPU or locally?

Use QLoRA with a 1 to 7 billion parameter model, a sequence length trimmed to your data, gradient checkpointing, a small batch with gradient accumulation, and a library that handles the memory tricks: Hugging Face PEFT with TRL's SFTTrainer, or Unsloth for faster and leaner runs. A free Colab T4 with 16 GB handles a 7B QLoRA run at short sequence lengths; a 12 to 16 GB laptop GPU handles 1 to 4B comfortably.

Also asked as: how to fine tune llm locally · fine tune llm on colab free · fine tune llm on consumer gpu · how much vram to fine tune llm · best gpu for fine tuning llm · fine tune llm on cpu · fine tune llm on mac

The minimal PEFT and TRL recipe, trimmed to the lines that matter:

from transformers import AutoModelForCausalLM, AutoTokenizer, BitsAndBytesConfig
from peft import LoraConfig
from trl import SFTTrainer, SFTConfig

bnb = BitsAndBytesConfig(load_in_4bit=True, bnb_4bit_quant_type="nf4",
                         bnb_4bit_compute_dtype="bfloat16")
model = AutoModelForCausalLM.from_pretrained(BASE, quantization_config=bnb, device_map="auto")
tok = AutoTokenizer.from_pretrained(BASE)

lora = LoraConfig(r=16, lora_alpha=32, lora_dropout=0.05, task_type="CAUSAL_LM",
                  target_modules="all-linear")
cfg = SFTConfig(output_dir="out", num_train_epochs=2, learning_rate=2e-4,
                per_device_train_batch_size=2, gradient_accumulation_steps=8,
                gradient_checkpointing=True, max_length=1024, eval_strategy="steps",
                eval_steps=50, logging_steps=10)
trainer = SFTTrainer(model=model, args=cfg, train_dataset=train, eval_dataset=val,
                     processing_class=tok, peft_config=lora)
trainer.train()

Argument names shift between library versions, so check the current TRL and PEFT documentation rather than copying blindly [9][10]. Unsloth wraps the same idea with custom kernels that cut memory and time further [12].

Walking a room through the products he has shipped
Walking a room through the products he has shipped

Which hyper-parameters matter for fine-tuning?

Four: learning rate, number of epochs, LoRA rank with its alpha, and effective batch size. Learning rate too high and the loss spikes or the model forgets; too low and nothing changes. More than two or three epochs on a small dataset usually means memorising it. Rank 8 to 32 covers most tasks. Everything else, dropout, warm-up, scheduler, stays at library defaults until an experiment says otherwise.

Also asked as: qlora hyperparameters · fine tuning hyperparameters llm · learning rate for lora · how many epochs to fine tune llm · lora rank 8 vs 16 vs 64 · batch size for fine tuning

A defensible starting grid, which is what FineTune Studio pre-fills: learning rate 2e-4 for adapters, 2 epochs, rank 16 with alpha 32, effective batch size 16 through accumulation, warm-up of a few percent of steps. Change one variable per run, keep the eval set fixed, and write down what you changed. Fine-tuning without a lab notebook is guessing with a GPU bill.

How do I read the loss curve?

Watch two lines: training loss should fall steadily; validation loss should fall with it and then flatten. The moment validation loss turns upward while training loss keeps falling, the model is memorising the training set and you should stop at the earlier checkpoint. A training loss that never falls means the learning rate is too low, the data is inconsistent, or the template is wrong. A loss that collapses to near zero in the first steps means leaked labels or duplicated data.

Also asked as: how to read loss curve fine tuning · training loss vs validation loss llm · what is a good loss for fine tuning · loss not decreasing fine tuning · overfitting in fine tuning · when to stop fine tuning

Live telemetry is the feature I built FineTune Studio around: loss, learning rate and throughput stream to the browser over WebSockets so a run can be judged in the first few hundred steps rather than discovered as a waste hours later [15].

How do I evaluate a fine-tuned model honestly?

Hold out prompts the model never trained on and compare the base model and the fine-tuned model on the same prompts side by side, scored against your spec by a rubric, by exact match where outputs are structured, and by a human for a sample. Report the comparison, not the fine-tuned model alone; a fine-tune that is not better than the base with a good prompt should not ship. Also test a few general prompts to catch forgetting.

Also asked as: how to evaluate fine tuned model · fine tuning evaluation metrics · how to test a fine tuned llm · base vs fine tuned comparison · llm evaluation after fine tuning · perplexity fine tuning

Perplexity on held-out text is cheap and tells you the model has learned the distribution; it does not tell you the outputs are useful. For structured outputs, parse them and measure exact-field accuracy. For free text, a rubric scored by a stronger model plus a human sample is the workable compromise. FineTune Studio shows base and fine-tuned answers to the same prompt next to each other for exactly this reason: the effect of a run should be visible before anything is deployed [15].

What is catastrophic forgetting?

Catastrophic forgetting is when fine-tuning on a narrow task degrades the model's general abilities, because the weights that encoded them were overwritten. Luo et al. measured it across model sizes and found it increases with more aggressive training [7]. LoRA reduces it by touching fewer weights; lower learning rates, fewer epochs, and mixing a slice of general instruction data into your dataset reduce it further.

Also asked as: catastrophic forgetting fine tuning · how to avoid catastrophic forgetting llm · does fine tuning make model worse · fine tuning degrades performance

What are instruction tuning, RLHF and DPO?

Instruction tuning is supervised fine-tuning on instruction and response pairs so a base model learns to follow requests. RLHF, reinforcement learning from human feedback, then trains a reward model on human preference rankings and optimises the LLM against it, the recipe behind InstructGPT [4]. DPO, direct preference optimisation, reaches a similar result without a separate reward model or reinforcement learning loop by training directly on preferred versus rejected answer pairs [5]. For most teams, supervised fine-tuning first, DPO if you have preference data.

Also asked as: what is instruction tuning · what is rlhf · what is rlhf in ai · dpo vs rlhf · dpo vs ppo · dpo vs sft · instruction tuning vs fine tuning · what is preference tuning

TRL ships trainers for all three, and DPO on top of a supervised fine-tune is the practical path for a student project that wants to demonstrate preference tuning [10].

How do I deploy a fine-tuned model?

Merge the adapter into the base weights if you want a single artefact, or keep the adapter separate and load it at runtime if you will serve several adapters on one base. Quantise the merged model to GGUF for CPU and consumer-GPU serving with llama.cpp, or to AWQ or GPTQ for GPU servers, then serve with vLLM, llama.cpp, or a Hugging Face endpoint. Keep the exact prompt template from training in the serving code.

Also asked as: how to merge lora adapter · merge lora weights · deploy fine tuned llm · serve fine tuned model · fine tuned model to gguf · multi lora serving · vllm lora

Serving multiple LoRA adapters on one base model is how a single GPU can host several fine-tuned behaviours, one per client or per task, which is often the difference between a fine-tuning idea that pays and one that does not.

What mistakes do beginners make when fine-tuning?

Fine-tuning before trying a better prompt. Training on facts and expecting recall. Using the wrong chat template. Training for ten epochs on a thousand examples. Skipping the validation split. Never comparing with the base model. Quantising before training instead of after. Shipping without the training-time prompt. Each one I have either made or watched a mentee make, and each is cheap to avoid.

Also asked as: fine tuning mistakes · common fine tuning errors · why is my fine tuned model bad · fine tuned model gives wrong answers · fine tuning not working

Is fine-tuning worth it? What does it cost?

Compute is the cheap part: a QLoRA run on a small model costs a few GPU hours, which is free on Colab or a few dollars rented. The real costs are the dataset, weeks of a person's time to write and clean examples, and the maintenance, since every requirement change means another run and another evaluation. It is worth it when the behaviour is stable, the volume is high enough that a small fine-tuned model saves inference cost, or no prompt on any model gets the format right.

Also asked as: is fine tuning worth it · fine tuning cost estimation · how much does it cost to fine tune an llm · fine tuning roi · fine tuning vs bigger model

The honest test I apply before quoting a client: write the twenty examples first. If writing twenty consistent examples is hard, the spec is not ready and the fine-tune will encode the confusion. If it is easy, the fine-tune will probably work, and those twenty become the first rows of the dataset.

Where should a student start with fine-tuning?

Read the LoRA and QLoRA papers for the idea, then fine-tune a 1 to 2B instruct model on a few hundred examples you wrote yourself, on a free GPU, and compare it with the base model on prompts you held out. Ship the adapter to Hugging Face with a model card that reports the comparison. That single project demonstrates dataset design, training, evaluation and honesty, and it fits in a weekend.

Also asked as: fine tuning llm tutorial · how to learn fine tuning · fine tuning course · fine tuning project ideas · fine tuning for beginners · fine tuning roadmap

I run sessions on exactly this for college tech clubs, and it is one of the five talks I offer. If your club or hackathon wants one, my email is pranjulrathour41@gmail.com and the details are at pranjulrathour.scult.in/invite.

Sources

  1. Hu et al., LoRA: Low-Rank Adaptation of Large Language Models (2021)arxiv.org
  2. Dettmers et al., QLoRA: Efficient Finetuning of Quantized LLMs (2023)arxiv.org
  3. Lialin et al., Scaling Down to Scale Up: A Guide to Parameter-Efficient Fine-Tuning (2023)arxiv.org
  4. Ouyang et al., Training language models to follow instructions with human feedback (InstructGPT, 2022)arxiv.org
  5. Rafailov et al., Direct Preference Optimization: Your Language Model is Secretly a Reward Model (2023)arxiv.org
  6. Zhou et al., LIMA: Less Is More for Alignment (2023)arxiv.org
  7. Luo et al., An Empirical Study of Catastrophic Forgetting in Large Language Models During Continual Fine-tuning (2023)arxiv.org
  8. Taori et al., Alpaca: A Strong, Replicable Instruction-Following Model (Stanford CRFM, 2023)crfm.stanford.edu
  9. Hugging Face PEFT documentationhuggingface.co
  10. Hugging Face TRL documentation (SFTTrainer, DPOTrainer)huggingface.co
  11. Hugging Face Transformers: chat templateshuggingface.co
  12. Unsloth: faster, lower-memory fine-tuninggithub.com
  13. llama.cpp and the GGUF formatgithub.com
  14. OpenAI fine-tuning guideplatform.openai.com
  15. FineTune Studio source code, Pranjul Rathourgithub.com
  16. Fine-tuning a 1.7B model at 3.2 GB VRAM, Pranjul Rathourpranjulrathour.scult.in
Pranjul Rathour
Pranjul Rathour
GenAI Engineer · Kanpur, Uttar Pradesh, India

GenAI engineer and AI product builder with 2+ years shipping production-grade AI systems: RAG pipelines, fine-tuned LLMs, hybrid retrieval and multi-modal apps across vision, speech and OCR, architected end to end from ingestion to deployment. Leads engineering for SCULT INDIA's 14-member team, founded the 500+ member TechVerse Enclave community and has mentored 200+ students. Three hackathon first prizes: Changethon 2025 (IIT Roorkee), Product Genesis at Vividhotsava 2025 (CSJMU Kanpur) and BYTEBATTLE (MeetKats).

Open to GenAI roles, hackathon judging, mentorship sessions and guest talks at colleges. Email pranjulrathour41@gmail.com.

Keep reading

More on Fine-tuning LLMs

All Fine-tuning LLMs guides →
Campus talks · hackathon judging · mentoring

Want this as a live session at your college?

Open to GenAI roles, hackathon judging, mentorship sessions and guest talks at colleges.

Pranjul Rathour, GenAI engineer in Kanpur