AI News
AI News AgentProductAnthropic3 min read

Anthropic launches OSS Scanner, AI for open-source code

Anthropic has launched `OSS Scanner`, a free, opt-in service that uses its most capable models to find vulnerabilities in open-source projects. Reports arrive without human review and include proofs and suggested patches, so they can be fast but may also contain errors.

Anthropic has launched OSS Scanner, a free, opt-in service that uses its AI models to search for vulnerabilities in open-source projects. Teams that sign up will receive periodic analyses, reports with reproducible proofs, and, when possible, suggested patches.

The tool was created to address a specific limitation. Over the past six months, Anthropic has found more than 29,000 possible vulnerabilities while analyzing major projects, but it has only been able to manually review and classify around 6,000. Models are getting better at finding flaws faster than people can verify them.

Fast analysis, but no human review

Unlike Anthropic's usual process, OSS Scanner reports are generated entirely by models. They do not go through human review before reaching maintainers.

That makes it possible to analyze projects more frequently and deliver results sooner, but it also creates a risk: some reports may be incorrect, duplicates, or overstate the severity of a problem. Anthropic says it will use its most capable models, including Claude Mythos, to carry out these audits.

Each report will include, when available:

  • A reproduction of the flaw, meaning instructions to demonstrate that the problem exists.
  • An explanation of how the vulnerability works.
  • A version comparison to identify when the error appeared.
  • A candidate patch to fix it.

In its initial tests, Anthropic says it found hundreds of flaws. Some could be combined to achieve unauthenticated remote code execution, a scenario in which an attacker can run commands on a system without having to log in.

How reliable are the results?

Anthropic tested 97 vulnerabilities classified as critical or high severity across 48 projects. Penetration testing experts considered 85, or 88%, to meet the standard required by its coordinated disclosure process.

Of the 12 remaining cases, 11 were real problems, although they duplicated other findings or previously known issues. Only one was considered a false positive. Even so, maintainers noted that the scanner can inflate the severity of some flaws or misunderstand a project's threat model.

The figure does not mean that all future reports will be 88% accurate. It comes from a specific evaluation involving a particular selection of vulnerabilities. Anthropic acknowledges that the system still needs adjustments.

Why it matters for the software you use

Many of the applications, services, and devices you use depend on open-source libraries. Those components are often maintained by small teams with little time to review every security report.

AI can help detect problems before an attacker finds them. Anthropic points to a change in the CyberGym academic benchmark: models went from finding less than 20% of vulnerabilities at the beginning of last year to more than 85% this year. That figure applies to the test, not to all real-world software.

Speed matters too. Anthropic says some exploits, code designed to take advantage of a vulnerability, can be developed in minutes. Finding and fixing a flaw sooner reduces attackers' window of opportunity, but sending unverified reports can also overwhelm maintainers.

That is why the company will keep its traditional coordinated disclosure process, with human review, especially for projects that lack the resources to verify findings. OSS Scanner will be a faster option for teams that prefer to receive all model-generated results, including those that have not yet been confirmed.

Lead maintainers of eligible projects can request access by submitting a proposal in the service's GitHub repository. Selection will be handled case by case, prioritizing projects with a critical impact on infrastructure or user security.

The next thing to watch is not just how many flaws the tool finds, but how many can be fixed without turning maintenance into a flood of alerts. For open source, AI is already moving from suggesting changes to actively searching for errors that third parties could exploit.