RLHF

Reinforcement learning from human feedback. Human raters compare pairs of model outputs, a reward model learns to predict which one they prefer, and the instruction-tuned model is then optimized against that reward signal. It shapes tone, helpfulness, and refusal behavior after instruction tuning has already taught the model to follow a request.

Why exams ask this

Tested as the step after instruction tuning, not a replacement for it. The exam wants the sequence named: pretrain, then instruction-tune, then align preferences with human feedback, and flags an answer that skips straight from pretraining to RLHF.

Relevant to

Related concepts

Resources

No resources linked to this concept yet.