Safety · Fast-moving · Advanced
AI Alignment
Making an AI system pursue the goals its developers and users actually intend, including implicit expectations they never stated.
What AI Alignment is
Alignment splits into outer alignment — specifying the right objective — and inner alignment — ensuring the trained system actually adopts it rather than a correlated proxy.
How it works
Current practice relies on instruction tuning, preference optimisation, constitutional methods, evaluation against behavioural specifications and interpretability research to inspect internals.
Why it matters
As systems take more autonomous action, the gap between what we asked for and what we meant becomes an operational risk rather than a philosophical one.
Common uses
- →Assistant behaviour specification
- →Refusal and safety policy
- →Agent objective design
Strengths
- ✓Directly improves everyday usefulness as well as safety
Watch for
- ✓Whose values is a contested question
- ✓No verification method for general behaviour
Continue exploring
More in this collection
Browse all AI Concepts