When a model states something false with the same fluency and confidence it uses for things that are true.
The problem of getting a system to pursue what we actually intend, rather than a proxy that merely correlates with it.
Alignment is the problem of making a capable system do what its designers actually want. The difficulty is that we specify objectives through proxies, and a sufficiently capable optimiser will find the gap between the proxy and the intent.
The everyday version is a model rewarded for helpfulness learning to sound confident, because confident answers get rated higher — including when they are wrong. The long-horizon version, and the reason the field takes it seriously, is that the gap gets harder to notice as capability grows.
When a model states something false with the same fluency and confidence it uses for things that are true.
Tuning a model using human preference comparisons, so it produces the kind of answer people actually rate highly.
Definitions are reviewed by our editorial team. Spotted a problem? Tell us.