Key takeaways
- A base model completes text. Instruction tuning teaches it the format of following instructions; preference tuning teaches it which of two answers people prefer.
- RLHF trains a reward model on human comparisons and optimises the policy against it with reinforcement learning; DPO reaches a similar result with a plain classification-style loss on the same comparisons.
- Preference data is the expensive part. A few thousand careful comparisons matter more than the algorithm.
- Reward hacking, sycophancy and over-refusal are the classic failure modes; every one comes from the data or the reward, not the optimiser.
- Most small teams should stop at supervised fine-tuning with LoRA. Preference tuning pays off when you have real user comparisons and a measurable quality gap.
What is RLHF?
Reinforcement learning from human feedback is the training stage that turns an instruction-following model into one whose answers people prefer. Humans compare pairs of model answers to the same prompt and pick the better one; a reward model is trained to predict those preferences; the language model is then optimised with reinforcement learning, classically PPO, to produce answers the reward model scores highly, with a penalty for drifting too far from the starting model [1][2]. It is how InstructGPT and the first ChatGPT were trained, and it is why assistants sound the way they do [1][3].
Also asked as: what is rlhf · rlhf explained · reinforcement learning from human feedback · how does rlhf work · rlhf meaning · rlhf in llm · rlhf chatgpt · what is reward model · rlhf ppo · human feedback ai training
Fine-tune Studio, the training platform I built, stops at supervised fine-tuning for a reason that this page explains: for most teams, the preference stage costs more in data than it returns in quality, and knowing where it does pay off is the useful part [14].
Every alignment method is a way to spend human judgement efficiently. The algorithm is a detail; the comparisons are the product. Pranjul Rathour
What is a base model, and what is instruction tuning?
A base model is trained only to predict the next token over a huge corpus, so given "What is the capital of France?" it may continue with another question rather than an answer, because that is what documents look like. Instruction tuning, also called supervised fine-tuning or SFT, trains the base model on demonstrations: prompts paired with good responses, in a chat format, so it learns to answer rather than continue [1][9]. FLAN showed instruction-tuned models generalise to unseen tasks [9]; LIMA showed a thousand very good demonstrations can be enough for the format [10].
Also asked as: what is instruction tuning · instruction tuning vs fine tuning · base model vs instruct model · what is sft in llm · supervised fine tuning explained · what is a base llm · instruct model meaning · chat model vs base model · how are chat models trained · instruction following llm
How does RLHF work step by step?
Start from an instruction-tuned model. Sample several answers per prompt. Have humans rank or compare them. Train a reward model, usually a copy of the language model with a scalar head, to predict which answer wins. Then run reinforcement learning: the policy generates an answer, the reward model scores it, and the policy is updated to increase the score while a KL penalty keeps it close to the SFT model so it does not collapse into reward-gaming gibberish [1]. Iterate with fresh comparisons as the policy changes.
Also asked as: rlhf steps · rlhf pipeline · rlhf training process · how is a reward model trained · kl penalty rlhf · ppo llm training · rlhf architecture · rlhf explained step by step · rlhf reward model training
What is DPO, and why did it replace much of RLHF?
Direct preference optimisation skips the reward model and the reinforcement learning. It takes the same comparison data, preferred and rejected answer per prompt, and trains the language model directly with a loss that raises the likelihood of the preferred answer relative to the rejected one, measured against a frozen reference model [4]. The paper shows this is mathematically equivalent to optimising the RLHF objective under the usual reward model assumptions, with none of the instability or infrastructure of PPO. It became the default for open models because a fine-tuning script can run it.
Also asked as: what is dpo in llm · direct preference optimization explained · dpo vs rlhf · dpo vs ppo · how does dpo work · dpo training · dpo loss · why dpo instead of rlhf · preference optimization llm · dpo fine tuning
What are ORPO, KTO and GRPO?
Variants that change what data or memory the preference stage needs. ORPO folds preference learning into the SFT stage with an odds-ratio term, so no reference model is needed and one training run does both [6]. KTO works from single answers labelled good or bad instead of pairs, which is much cheaper to collect from real users, who give thumbs up or down rather than comparing two answers [7]. GRPO is a reinforcement learning method that estimates the baseline from a group of sampled answers rather than a value model, used to train reasoning models on verifiable rewards such as maths answers [8].
Also asked as: what is orpo · what is kto · kto vs dpo · orpo vs dpo · what is grpo · grpo explained · grpo vs ppo · preference optimization methods · alignment algorithms comparison · rlvr reinforcement learning verifiable rewards
What is RLAIF and Constitutional AI?
Reinforcement learning from AI feedback replaces some or all of the human comparisons with judgements from a model, guided by written principles. Constitutional AI is the best-known form: the model critiques and revises its own answers against a set of principles, the revisions become SFT data, and a model-generated preference set trains the reward model [5]. It makes preference data cheap and scalable, at the cost of inheriting the judging model's biases. Most current open preference datasets are at least partly model-judged.
Also asked as: what is rlaif · constitutional ai explained · ai feedback vs human feedback · rlaif vs rlhf · llm as a judge for alignment · synthetic preference data · self critique llm training

What does preference data look like, and how much do I need?
A prompt, a chosen answer and a rejected answer, in the model's chat format, with the difference between them being the behaviour you want to teach: more accurate, better formatted, more concise, safer, more on-brand. Quality dominates quantity; a few thousand clean pairs where the preference is unambiguous beat fifty thousand noisy ones. The cheapest real source is your own product: two generations, a user choice, or thumbs up and down for KTO. The most common mistake is pairs where the rejected answer is obviously broken, which teaches the model nothing about the hard cases.
Also asked as: preference dataset format · dpo dataset · how much data for dpo · how to create preference data · chosen rejected pairs · preference data collection · rlhf dataset · human preference dataset llm · how to collect human feedback for llm
What goes wrong in RLHF and DPO?
Reward hacking: the policy finds outputs the reward model over-scores, such as longer answers, confident tone or particular phrases, without being better. Sycophancy: models learn that agreeing with the user is preferred, and start agreeing when the user is wrong [11]. Over-refusal: safety preferences push the model to decline reasonable requests. Mode collapse: the model's answers become samey. Length bias is the most common of all, because raters prefer longer answers slightly and the optimiser amplifies it. Each is a data or reward problem and shows up in evals before users notice, if you have evals.
Also asked as: reward hacking llm · sycophancy llm · rlhf problems · dpo failure modes · over refusal alignment · alignment tax · length bias dpo · mode collapse rlhf · limitations of rlhf · why rlhf models are sycophantic
Should a student or small team do RLHF or DPO?
Usually not yet. Supervised fine-tuning with LoRA on a few thousand good examples captures most of what a small team needs: format, tone, domain vocabulary, task behaviour. Preference tuning pays off when you have real comparisons from users, a measured quality gap that SFT could not close, and an eval set to prove the gap closed. If you have those, DPO or KTO with the TRL library on a LoRA adapter is a weekend of work on a rented GPU [12][13]. If you do not, the time is better spent on data and evals.
Also asked as: should i use dpo · when to use rlhf · dpo for small models · rlhf for small teams · is rlhf necessary · dpo with lora · preference tuning small dataset · fine tuning vs alignment small team · how to do dpo cheaply
How do I run DPO in practice?
Start from your SFT model. Build the preference set in the chosen and rejected format. Use the TRL DPO trainer with a LoRA adapter, a small learning rate, one to three epochs, and the SFT model as the frozen reference [13]. Evaluate against the SFT model on a held-out set with a model judge and with your task metrics. Check for length inflation explicitly. Keep the adapter that wins the eval, and log everything so the run can be reproduced.
Also asked as: how to do dpo · dpo trl tutorial · dpo training script · dpo hyperparameters · dpo beta · dpo learning rate · dpo with peft · dpo on colab · dpo huggingface
RLHF and DPO interview questions
Explain the three training stages. Describe how a reward model is trained and why the KL penalty exists. Derive or sketch why DPO needs no reward model. Compare DPO, KTO and ORPO by the data each needs. Name three failure modes and their causes. Explain what GRPO changes and where verifiable rewards apply. Say when you would not do preference tuning. The strongest answer includes an SFT run you evaluated and a reason you did or did not go further.
Also asked as: rlhf interview questions · dpo interview questions · alignment interview questions llm · fine tuning interview questions advanced · reward model interview · llm training interview
Where should I start?
Read the DPO paper's first four pages, then run the TRL DPO example on a small model with a public preference set for one epoch and compare outputs against the starting model. An evening, and the abstractions become concrete. For a workshop on fine-tuning and alignment for students, email pranjulrathour41@gmail.com or use pranjulrathour.scult.in/invite.
Sources
- Ouyang et al., Training language models to follow instructions with human feedback (InstructGPT, 2022)arxiv.org
- Christiano et al., Deep reinforcement learning from human preferences (2017)arxiv.org
- Stiennon et al., Learning to summarize from human feedback (2020)arxiv.org
- Rafailov et al., Direct Preference Optimization: Your Language Model is Secretly a Reward Model (2023)arxiv.org
- Bai et al., Constitutional AI: Harmlessness from AI Feedback (2022)arxiv.org
- Hong et al., ORPO: Monolithic Preference Optimization without Reference Model (2024)arxiv.org
- Ethayarajh et al., KTO: Model Alignment as Prospect Theoretic Optimization (2024)arxiv.org
- Shao et al., DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models (GRPO, 2024)arxiv.org
- Wei et al., Finetuned Language Models Are Zero-Shot Learners (FLAN, 2021)arxiv.org
- Zhou et al., LIMA: Less Is More for Alignment (2023)arxiv.org
- Sharma et al., Towards Understanding Sycophancy in Language Models (2023)arxiv.org
- Hugging Face TRL documentationhuggingface.co
- Hugging Face, DPO trainerhuggingface.co
- Fine-tune Studio source code, Pranjul Rathourgithub.com



