Alignment

The problem of getting a system to pursue what we actually intend, rather than a proxy that merely correlates with it.

Ethics & Safety Intermediate 1 min read

Definition

Alignment is the problem of making a capable system do what its designers actually want. The difficulty is that we specify objectives through proxies, and a sufficiently capable optimiser will find the gap between the proxy and the intent.

The everyday version is a model rewarded for helpfulness learning to sound confident, because confident answers get rated higher — including when they are wrong. The long-horizon version, and the reason the field takes it seriously, is that the gap gets harder to notice as capability grows.

How it relates to other terms

See also RLHF Tuning a model using human preference comparisons, so it produces the kind of answer people actually rate high... See also Hallucination When a model states something false with the same fluency and confidence it uses for things that are true.
Referenced by

Related terms

When a model states something false with the same fluency and confidence it uses for things that are true.

Beginner Ethics & Safety

Tuning a model using human preference comparisons, so it produces the kind of answer people actually rate highly.

Advanced Ethics & Safety

Was this definition helpful?

Definitions are reviewed by our editorial team. Spotted a problem? Tell us.