Also called: reinforcement learning from human feedback
Training a model on human preferences between candidate answers, so it learns which responses people actually want.
Why it mattersThe step that turned raw text predictors into assistants worth talking to.
See also Alignment, Fine-Tuning