← World Of AI

Generative AI

RLHF

Reinforcement learning from human feedback — how a raw model becomes a usable assistant.

People rank alternative responses, a reward model is trained to predict those rankings, and the language model is then optimised against that reward.

It is what converts a next-token predictor into something that answers the question asked. It also inherits the taste and the blind spots of whoever produced the rankings, which is why annotation guidelines are a serious artefact.

Also in Generative AI

Apply

Begin the first module

Become AI native, it is the real deal today, and if it is not for you, you have lost nothing but learnt a new skill.