RLHF
Reinforcement Learning from Human Feedback
RLHF improves an AI model using feedback about which responses people prefer.
Explained simply
RLHF improves an AI model using feedback about which responses people prefer.
1AI responses
→
2Human preference
→
3Train from feedback
→
4Better behaviour
At a glance
- Category
- Models
- Difficulty
- Advanced
- Introduced
- 2017
Real example
Reviewers choose the better of two AI responses, helping the model learn which style people prefer.
Why it matters
It gives you a clear way to understand where this idea fits in AI and how it affects the tools you use.
Timeline
Origins
The ideas behind RLHF begin developing.
2017
RLHF becomes a recognised term or technique.
Wider use
Research and practical applications increase.
Modern AI
RLHF becomes connected to newer AI systems and products.
Today
RLHF remains relevant in the models area of AI.
Learn next
Continue with these connected terms:
Simple infographic
A quick visual way to understand RLHF.
1AI responses
→
2Human preference
→
3Train from feedback
→
4Better behaviour
1 of 6
Was this helpful?
Your feedback helps improve this explanation.
