How to prepare a dataset for fine-tuning an LLM: formats, chat templates, how much data, synthetic data, validation and the checklist before you train

The dataset is the fine-tune. How to write a spec, choose between Alpaca, ShareGPT and chat formats, apply the model's chat template correctly, decide how many examples you need, generate and filter synthetic data, deduplicate, split, and validate every row before spending a GPU hour. From building a fine-tuning platform.

Requirements gathering and user flows, on stage
Requirements gathering and user flows, on stage

Key takeaways

  • Write a one-paragraph spec of the behaviour first. Every example must follow it exactly; inconsistency is what the model learns.
  • Format is JSON; what matters is the chat template the tokenizer renders. Use the tokenizer's template function, never hand-write special tokens.
  • Hundreds of clean examples change style and format; thousands teach a skill; tens of thousands adapt to a domain. Below a hundred, prompt instead.
  • Synthetic data works when filtered hard against the spec and mixed with real examples; unfiltered, it teaches the generator's habits.
  • Deduplicate, hold out ten to twenty percent, measure lengths, mask labels to the assistant turn, and read fifty rows by eye before training.

How do I prepare a dataset for fine-tuning an LLM?

Start with a written spec of the behaviour you want, then collect or write examples that follow it exactly: for each, the input a user would give and the ideal output, in the chat format the base model expects. Deduplicate, validate every row against the spec, hold out a validation split, measure token lengths, and read a random sample by eye. The dataset is the fine-tune; the training run only copies what is in it, including its inconsistencies.

Also asked as: how to prepare dataset for fine tuning · how to create dataset for llm fine tuning · fine tuning dataset preparation · dataset for fine tuning llm · how to make training data for llm · fine tuning data format · llm training data preparation · how to build a fine tuning dataset

FineTune Studio, the platform I built in 2026, validates the uploaded dataset before it lets a run start, because the most common reason a run produced a bad model had nothing to do with hyper-parameters: the data was inconsistent, mis-templated or leaked into the validation split [12]. This page is that validation, explained.

Write twenty examples by hand before anything else. If twenty consistent examples are hard to write, the spec is not ready, and the fine-tune will faithfully learn the confusion. Pranjul Rathour, from building FineTune Studio

What should the spec contain?

One paragraph: who the user is, what they send, what the model must produce, in what format, with what tone, and what it must do when the input is unclear or out of scope. Then three example pairs that follow it. Every later example is checked against this paragraph. A spec you cannot write in a paragraph is a task you are not ready to train.

Also asked as: fine tuning task specification · how to define fine tuning task · what behavior to fine tune · fine tuning goals · fine tuning requirements document

Which dataset format should I use: Alpaca, ShareGPT or chat messages?

Use the messages format, a list of role and content turns, system, user, assistant, because it maps directly onto the chat template every modern model uses and onto what TRL's trainer expects [7]. Alpaca's instruction, input and output fields [1] and ShareGPT's multi-turn conversations are older JSON shapes that convert to messages trivially. The format is a container; what the model actually sees is the template rendered from it.

Also asked as: alpaca format vs sharegpt format · alpaca dataset format · sharegpt format · chatml format · fine tuning dataset json format · jsonl format for fine tuning · openai fine tuning format · conversational dataset format · instruction dataset format

A row in the messages format, which is also the shape most APIs accept [9]:

{"messages": [
  {"role": "system", "content": "You are a support agent for Acme. Answer only from the provided policy. If the policy does not cover the question, say so and offer to escalate."},
  {"role": "user", "content": "Policy: Refunds within 14 days with receipt.\n\nQuestion: I bought it 20 days ago, can I return it?"},
  {"role": "assistant", "content": "The policy allows refunds within 14 days of purchase with a receipt, so a return at 20 days is outside the window. I can escalate this to a supervisor if you would like."}
]}

What is a chat template and why does it matter?

A chat template is the exact string layout a model was trained on for conversations: the special tokens that mark the start and end of each turn and which role speaks. Every model family uses a different one. If you train with the wrong template, or hand-write the tokens with a small mistake, the model learns garbage or produces it at inference. Always render examples with the tokenizer's own template function [6], and use the same template when you serve the model.

Also asked as: chat template fine tuning · what is chat template · apply chat template hugging face · chatml vs llama template · special tokens fine tuning · eos token fine tuning · template mismatch fine tuning · how to format prompts for fine tuning

from transformers import AutoTokenizer
tok = AutoTokenizer.from_pretrained(BASE)
text = tok.apply_chat_template(row["messages"], tokenize=False, add_generation_prompt=False)
# never build this string by hand; the template lives with the model

Template mismatch is the single most common cause of a fine-tuned chat model that "does not work", in my experience and in the support questions FineTune Studio users asked. My longer note on formats and templates is on the portfolio [11].

How much data do I need to fine-tune an LLM?

For tone, style and output format: a few hundred high-quality examples produce a visible, stable change. For a classification or extraction skill: a few thousand covering the classes and the edge cases. For adapting to a domain's language: tens of thousands. Below about a hundred, put the examples in the prompt instead. LIMA showed that a thousand carefully curated examples were enough to teach strong instruction-following, and that quality mattered more than quantity [2].

Also asked as: how much data to fine tune llm · minimum dataset size for fine tuning · how many examples for fine tuning · fine tuning with small dataset · dataset size for lora · how many samples for qlora · fine tuning with 100 examples · fine tuning with 1000 examples

These are orders of magnitude from experience, not laws; the only reliable number comes from training on half your data and checking whether the other half still improves the model.

Can I use synthetic data for fine-tuning?

Yes, and most fine-tuning datasets today are partly synthetic: a stronger model generates examples from your spec, from seed examples, or from documents. Self-Instruct and Alpaca showed the approach at scale [3][1]. It works when you filter hard: validate every generated row against the spec, remove duplicates and near-duplicates, spot-check by hand, and mix in real examples so the model does not learn the generator's tics and refusals. Unfiltered synthetic data teaches a copy of the generator, not your task.

Also asked as: synthetic data for fine tuning · how to generate training data with llm · synthetic dataset generation llm · is synthetic data good for fine tuning · self instruct · data augmentation for llm fine tuning · llm generated training data quality

How do I clean and deduplicate the dataset?

Normalise whitespace and encoding. Remove exact duplicates, then near-duplicates by comparing normalised text or embeddings, because duplicated training data both wastes compute and makes models memorise [5]. Remove rows that violate the spec: wrong format, wrong language, refusals you did not intend, outputs that reference the prompt. Strip personal data you do not have consent to train on. Measure token lengths and trim or split outliers so no row exceeds your training sequence length.

Also asked as: dataset cleaning for fine tuning · deduplicate training data · data quality for llm training · remove duplicates from dataset · dataset validation llm · training data quality checks · how to clean text data for fine tuning

The Hugging Face Datasets library handles the mechanics, loading, mapping, filtering and splitting, for datasets of any size [8]. My own pre-training validation checklist, which FineTune Studio runs automatically, is written up separately [10].

Requirements gathering and user flows, on stage
Requirements gathering and user flows, on stage

How do I split training and validation data?

Hold out ten to twenty percent, chosen so that near-duplicates of a training row never land in validation, or the validation loss lies to you. For grouped data, all rows from one document, one user or one conversation go to the same split. Keep a third, small test set of examples you wrote by hand and never look at during tuning, for the final honest comparison of base versus fine-tuned.

Also asked as: train validation split fine tuning · how to split dataset for llm · validation set for fine tuning · test set for llm evaluation · data leakage fine tuning · holdout set llm

What is label masking, and should I train on the user turns?

Label masking means computing the loss only on the assistant's tokens, so the model learns to produce answers rather than to predict the user's questions. Most trainers do this for you when you use the messages format with a chat template and an assistant-only loss option [7]. Training on user turns wastes capacity and can teach the model to continue user text. Mask them unless you have a specific reason not to.

Also asked as: label masking fine tuning · train on completions only · assistant only loss · loss on user tokens · completion only training · sft loss masking

What are common dataset mistakes?

Mixing two specs in one dataset, so the model learns to flip between them. Hand-writing chat tokens. Duplicated rows inflating confidence. A validation split that shares documents with training. Outputs that mention the instruction. Ten variations of the same easy example and none of the hard cases. Personal data nobody consented to. Not reading any rows. A model can only be as consistent as its data, and it will be exactly as inconsistent.

Also asked as: fine tuning dataset mistakes · common data mistakes llm training · why fine tuned model is bad data · bad training data examples · fine tuning data pitfalls

Dataset interview questions

How much data do you need for a style fine-tune versus a skill? What is a chat template and what happens if you get it wrong? How do you generate synthetic data without teaching the generator's habits? How do you prevent leakage between splits? What is label masking? Each answer above is a paragraph; the best evidence is a dataset you built and a model whose behaviour you can trace to specific rows.

Also asked as: fine tuning dataset interview questions · llm data preparation interview · training data interview questions

Where should I start?

Pick one narrow behaviour, write the paragraph spec and twenty examples tonight, render them with the tokenizer's template, and read them back. Then generate two hundred more synthetically, filter them against the spec, and hold out twenty. That dataset will fine-tune a small model on a free GPU tomorrow. For a hands-on fine-tuning workshop at your college, email pranjulrathour41@gmail.com or use pranjulrathour.scult.in/invite.

Sources

  1. Taori et al., Alpaca: A Strong, Replicable Instruction-Following Model (Stanford CRFM, 2023)crfm.stanford.edu
  2. Zhou et al., LIMA: Less Is More for Alignment (2023)arxiv.org
  3. Wang et al., Self-Instruct: Aligning Language Models with Self-Generated Instructions (2022)arxiv.org
  4. Ouyang et al., Training language models to follow instructions with human feedback (2022)arxiv.org
  5. Lee et al., Deduplicating Training Data Makes Language Models Better (2021)arxiv.org
  6. Hugging Face Transformers: chat templateshuggingface.co
  7. Hugging Face TRL: SFTTrainer dataset formatshuggingface.co
  8. Hugging Face Datasets documentationhuggingface.co
  9. OpenAI fine-tuning guide: preparing your datasetplatform.openai.com
  10. Dataset validation before training, Pranjul Rathourpranjulrathour.scult.in
  11. Fine-tuning dataset formats: Alpaca, ShareGPT and chat templates, Pranjul Rathourpranjulrathour.scult.in
  12. FineTune Studio source code, Pranjul Rathourgithub.com
Pranjul Rathour
Pranjul Rathour
GenAI Engineer · Kanpur, Uttar Pradesh, India

GenAI engineer and AI product builder with 2+ years shipping production-grade AI systems: RAG pipelines, fine-tuned LLMs, hybrid retrieval and multi-modal apps across vision, speech and OCR, architected end to end from ingestion to deployment. Leads engineering for SCULT INDIA's 14-member team, founded the 500+ member TechVerse Enclave community and has mentored 200+ students. Three hackathon first prizes: Changethon 2025 (IIT Roorkee), Product Genesis at Vividhotsava 2025 (CSJMU Kanpur) and BYTEBATTLE (MeetKats).

Open to GenAI roles, hackathon judging, mentorship sessions and guest talks at colleges. Email pranjulrathour41@gmail.com.

Keep reading

More on Fine-tuning LLMs

All Fine-tuning LLMs guides →
Campus talks · hackathon judging · mentoring

Want this as a live session at your college?

Open to GenAI roles, hackathon judging, mentorship sessions and guest talks at colleges.

Pranjul Rathour, GenAI engineer in Kanpur