Skip to content
CW101
OngoingOfficialConcept

Constitutional AI

The training method behind Claude's behaviour — a written set of principles the model is trained against, rather than behaviour shaped only by human preference labels.

Last verified August 21, 2026 · 2 sources

Constitutional AI is how Anthropic trains Claude’s behaviour: instead of relying only on human labels for what is harmful, the model is trained against an explicit written constitution — a set of principles it critiques and revises its own outputs against.

Why it matters practically, not just philosophically

Two consequences you can actually observe:

The principles are published. You can read what Claude is trained to weigh. That is unusual, and it makes model behaviour arguable from a document rather than from guesswork.

Behaviour is a stated design target, not an emergent accident. When Claude declines something, there is a written basis for it. That does not make every decision correct, but it makes it inspectable.

What it is not

  • Not the Usage Policy. The constitution shapes how the model behaves. The usage policy governs what you are permitted to do with it. Different documents, different enforcement.
  • Not a guarantee. System cards exist precisely because trained behaviour has to be measured, not assumed.
  • Not static. The constitution has been revised and republished.

What this page does not claim

This entry does not reproduce the constitution’s contents or summarise its principles — it is a substantive document that deserves reading in full rather than in paraphrase.

Sources

  1. 01Constitutional AI - Harmlessness from AI FeedbackOfficialanthropic.com
  2. 02Claude's ConstitutionOfficialanthropic.com

Page changelog

  1. VerifiedConfirmed as published Anthropic research and an ongoing published document.

Start typing to search every entity in the reference.