← Back
ModelsAdvanced2017

RLHF

Reinforcement Learning from Human Feedback

RLHF improves an AI model using feedback about which responses people prefer.

Explained simply

RLHF improves an AI model using feedback about which responses people prefer.

1AI responses
2Human preference
3Train from feedback
4Better behaviour

At a glance

Category
Models
Difficulty
Advanced
Introduced
2017

Real example

Reviewers choose the better of two AI responses, helping the model learn which style people prefer.

Why it matters

It gives you a clear way to understand where this idea fits in AI and how it affects the tools you use.

Timeline

Origins

The ideas behind RLHF begin developing.

2017

RLHF becomes a recognised term or technique.

Wider use

Research and practical applications increase.

Modern AI

RLHF becomes connected to newer AI systems and products.

Today

RLHF remains relevant in the models area of AI.

Dates for broad concepts may describe the period when the idea emerged or became widely used, rather than one exact invention date.

Learn next

Simple infographic

A quick visual way to understand RLHF.

1AI responses
2Human preference
3Train from feedback
4Better behaviour
1 of 6

Was this helpful?

Your feedback helps improve this explanation.