Alignment · Established · Advanced
Constitutional AI
An alignment approach where a model critiques and revises its own outputs against a written set of principles, reducing reliance on human harm labelling.
What Constitutional AI is
Instead of humans labelling every harmful response, an explicit constitution states the principles, and the model applies them to generate its own preference data.
How it works
The model produces a response, critiques it against the principles, revises it, and the revision pairs become training data for supervised and preference tuning.
Why it matters
It makes the value specification explicit and inspectable, which is easier to govern than values implicit in a crowd of annotators.
Common uses
- →Harm reduction training
- →Transparent behaviour policy
- →Scaling alignment data
Strengths
- ✓Written, auditable principles
- ✓Less exposure of humans to harmful content
Watch for
- ✓Principles still reflect their authors
- ✓Model must interpret them correctly
Continue exploring
More in this collection
Browse all AI ConceptsSources & References
Anthropic — Constitutional AI