← World Of AI
Generative AI
RLHF
Reinforcement learning from human feedback — how a raw model becomes a usable assistant.
People rank alternative responses, a reward model is trained to predict those rankings, and the language model is then optimised against that reward.
It is what converts a next-token predictor into something that answers the question asked. It also inherits the taste and the blind spots of whoever produced the rankings, which is why annotation guidelines are a serious artefact.
Apply
Begin the first module
Become AI native, it is the real deal today, and if it is not for you, you have lost nothing but learnt a new skill.