- Aligns model outputs with human preferences and values
- Improves helpfulness, honesty, and safety beyond raw pre-training
- Captures nuanced quality judgments hard to encode as rules
- Underpins the usefulness of modern conversational assistants
Humans rank model outputs, those rankings train a reward model, and reinforcement learning then optimises the model toward higher-rewarded responses.
It aligns models with nuanced human preferences for helpfulness and safety that are difficult to specify with explicit rules.
It depends on the quality and consistency of human feedback and can introduce the biases of its annotators.
Follow Techment on LinkedIn for practical AI, Data Engineering, and Microsoft Fabric insights delivered every week.
Hello popup window