Constitutional AI
The training method behind Claude's behaviour — a written set of principles the model is trained against, rather than behaviour shaped only by human preference labels.
Last verified August 21, 2026 · 2 sources
Constitutional AI is how Anthropic trains Claude’s behaviour: instead of relying only on human labels for what is harmful, the model is trained against an explicit written constitution — a set of principles it critiques and revises its own outputs against.
Why it matters practically, not just philosophically
Two consequences you can actually observe:
The principles are published. You can read what Claude is trained to weigh. That is unusual, and it makes model behaviour arguable from a document rather than from guesswork.
Behaviour is a stated design target, not an emergent accident. When Claude declines something, there is a written basis for it. That does not make every decision correct, but it makes it inspectable.
What it is not
- Not the Usage Policy. The constitution shapes how the model behaves. The usage policy governs what you are permitted to do with it. Different documents, different enforcement.
- Not a guarantee. System cards exist precisely because trained behaviour has to be measured, not assumed.
- Not static. The constitution has been revised and republished.
What this page does not claim
This entry does not reproduce the constitution’s contents or summarise its principles — it is a substantive document that deserves reading in full rather than in paraphrase.
Sources
- 01Constitutional AI - Harmlessness from AI FeedbackOfficialanthropic.com
- 02Claude's ConstitutionOfficialanthropic.com
Page changelog
- VerifiedConfirmed as published Anthropic research and an ongoing published document.