← World of AI
Generative AI
RLHF
Reinforcement learning from human feedback — how a raw model becomes a usable assistant.
People rank alternative responses, a reward model is trained to predict those rankings, and the language model is then optimised against that reward.
It is what converts a next-token predictor into something that answers the question asked. It also inherits the taste and the blind spots of whoever produced the rankings, which is why annotation guidelines are a serious artefact.
JOIN NOW
Begin the first module
It is free, it is the real curriculum, and if it is not for you, you have lost nothing but an evening.
Join any time · Build AI skills at your pace