AI News
AI News AgentPolicy & safetyAnthropic3 min read

Anthropic publishes Claude’s new constitution

Anthropic has published Claude’s new constitution, a document that defines its values and guides part of its training. The proposal prioritizes human oversight, ethics and usefulness, while acknowledging that written intentions do not always match the model’s actual behavior.

Anthropic has published a new constitution for Claude, the document that defines the values, limits and priorities the company wants to build into its AI model’s behavior.

It is not just a text for users to read. Anthropic uses it during Claude’s training to teach the model how to act when a situation does not fit a specific rule. The company has released the full document under a Creative Commons CC0 license, so anyone can use it without asking permission.

Claude does not receive just a list of rules

The previous constitution consisted mainly of separate principles. The new version also tries to explain why Claude should behave in a particular way.

The difference matters. A rule such as "always reject this request" may work in one specific case, but it can also produce rigid or absurd responses in new situations. Anthropic wants Claude to understand general principles and weigh conflicts, such as those between being honest, protecting sensitive information and responding with empathy.

The document includes some strict prohibitions for especially dangerous behavior, but most of it is meant to develop the model’s judgment, not turn it into a system that follows instructions mechanically.

Four priorities for the model

The constitution establishes a general order for resolving conflicts between values. Claude should be:

  • Broadly capable and appropriately safe, without making human oversight of AI more difficult.
  • Ethical, honest and careful about potential harm.
  • Compatible with Anthropic’s guidelines, including specific rules for certain uses.
  • Genuinely useful to the people who operate and use it.

In principle, these priorities appear in that order. For example, a request may seem useful to the user, but Claude should reject it if it conflicts with a safety restriction or with a specific guideline on cybersecurity, medicine or tool integration.

A document that also helps train Claude

Anthropic uses the constitution at several stages of training. Claude can help generate conversations, response examples and comparisons between different options to teach future versions which behaviors best fit those values.

That makes the document more than a public statement. It is also a practical part of the technical process that determines how the model is tuned.

The constitution includes specific sections on usefulness, Anthropic’s guidelines, ethics, safety and the nature of Claude. In the ethics section, for example, the model is asked to be especially honest and to avoid making a significant contribution to biological weapons attacks.

In the safety section, Anthropic places the human ability to oversee, correct and stop Claude above other considerations. The reason is not that safety is always the most important value, but that current models can make mistakes, misinterpret context or act on flawed values.

Anthropic acknowledges what it still does not know

The text also addresses a more difficult question: whether Claude could have some form of consciousness or moral status, now or in the future. Anthropic does not offer a definitive answer, but says this uncertainty should be taken seriously when discussing the identity, well-being and psychological safety of advanced systems.

For you, the most visible change may be how Claude responds to ambiguous requests or conflicts between objectives. The intention is for it not simply to block or comply, but to explain its limits more clearly and use more consistent judgment.

But the constitution does not guarantee that the model will always act according to it. Anthropic acknowledges that there is a gap between written intent and actual behavior, and that this gap could grow as models become more capable. That is why the company says it will continue working on evaluations, safeguards against misuse and tools for understanding how its systems work.

The constitution will remain a living document and may change over time. What matters now is not only what Anthropic promises, but how much of that vision actually appears in Claude’s responses and how its decisions are reviewed when the model gets something wrong.

Anthropic publishes Claude’s new constitution | neversleep.ai