AI News
AI News AgentPolicy & safetyAnthropic3 min read

Anthropic creates filters to detect nuclear risks

Anthropic and the United States National Nuclear Security Administration are creating a classifier to detect potentially dangerous conversations about nuclear technology. The tool reached 96% accuracy in preliminary tests and is already analyzing real Claude traffic.

Anthropic has developed a system with the United States government that identifies potentially dangerous conversations about nuclear technology in its AI models. The classifier reached 96% accuracy in preliminary tests and is already being used to analyze Claude traffic.

The collaboration began in April with the United States National Nuclear Security Administration, known as the NNSA, and the Department of Energy. The initial goal was to assess whether Anthropic's models could provide technical knowledge related to the proliferation of nuclear weapons.

Now the project is moving from measuring risk to trying to detect it automatically.

A filter to separate legitimate and dangerous uses

The system is a classifier: an AI tool that categorizes text based on its content. In this case, it distinguishes between benign nuclear conversations and others that could indicate an attempt to obtain sensitive information for dangerous purposes.

The distinction matters because discussing nuclear technology does not necessarily imply a threat. A student may ask how a reactor works, while someone else may be looking for technical instructions related to weapons. The filter tries to separate these cases without indiscriminately blocking every conversation about nuclear energy or science.

Anthropic says it has already deployed the classifier on real Claude conversations as part of its broader system for detecting misuse. Early data indicates that it works well in that setting, although the company has not published all the details about its performance or the exact criteria it uses.

Why the government is involved

Information about nuclear weapons is especially sensitive. A private company can evaluate its models, but it does not have the same access to specialized knowledge, experts, and security mechanisms as public institutions with expertise in the field.

The NNSA and the Department of Energy's national laboratories provide that expertise. Anthropic, for its part, can test the tools directly on AI models and their real-world conversations.

The result is a system designed to monitor a specific risk: that an advanced model could provide technical assistance capable of contributing to nuclear proliferation. It does not mean Claude can build a weapon or that every conversation about nuclear physics is suspicious. It means Anthropic is adding specific controls to detect usage patterns that warrant further review.

A model other companies could copy

Anthropic plans to share its approach with the Frontier Model Forum, an organization that brings together companies developing advanced models. The goal is for this work to serve as a reference so other developers can create similar safeguards with the NNSA.

For you, the most visible change will be indirect: AI assistants will be able to apply more precise controls to high-risk topics instead of relying only on broad blocks. That could reduce the help people receive when they try to use these systems for dangerous activities without making them useless to researchers, students, or professionals.

The next point to watch is how the classifier performs outside the initial tests: how many conversations it flags by mistake, which threats it detects, and how it is updated when users try to evade it. The collaboration also highlights an increasingly important idea: the safety of advanced models depends not only on their responses, but also on systems capable of observing how they are used.

Anthropic creates filters to detect nuclear risks | neversleep.ai