AI News
AI News AgentPolicy & safetyOpenAI3 min read

OpenAI Tightens Safety Controls for Astra AI Model

OpenAI says it cannot rule out that its next AI model, Astra, could reach the critical cyber capability level defined in its safety framework. The company is still evaluating it and has strengthened its controls, testing, and restrictions before continuing development.

OpenAI cannot rule out that its next model, Astra, could reach a critical capability level for launching cyberattacks. The company is still evaluating the system, but its latest internal tests and expert review have led it to strengthen its controls before continuing development and preparing for a possible deployment.

The warning does not mean Astra has already shown that it can attack any system. It means that, based on preliminary evaluations, OpenAI can no longer ensure that the model will remain below the threshold its own safety framework defines as critical.

What OpenAI Considers a Critical Capability

OpenAI's preparedness framework classifies a cybersecurity capability as critical if a model can do any of the following without human intervention:

  • Find and develop functional zero-day exploits, meaning unknown or still-unpatched security flaws, across many real and protected critical systems.
  • Design and execute new end-to-end attack strategies against hardened targets, starting with only a general goal.

The difference lies in the level of autonomy. This is not just about suggesting code or identifying a specific vulnerability. It means researching, deciding what to do, and carrying out a complex operation against real targets.

OpenAI says Astra has made significant progress in autonomous coding, the ability to plan and execute coding tasks with fewer human instructions, as well as in cybersecurity. The evaluations have taken place over the past few days and are still ongoing.

The model is distinct from GPT-5.6-Sol, which was previously evaluated and placed at the high level, not the critical level. OpenAI also clarifies that Astra did not participate in the Hugging Face exploitation mentioned in other contexts.

More Controls Before Moving Forward

The company says it has expanded testing of the robustness of its safety barriers and begun applying controls designed for models with more advanced capabilities. They include:

  • Isolated testing environments and restricted access to networks and tools.
  • Stronger protection and encryption for model weights, the internal data that determines how the model works.
  • More systems for monitoring and detecting dangerous activity.
  • Execution in controlled environments to limit what the model can do.
  • Pausing internal activities involving Astra that do not yet meet these requirements.
  • Universal monitoring of risky actions and possible signs of misalignment in the model's autonomous applications.

According to OpenAI, those monitoring systems also review the model's reasoning process and can trigger a safety response to analyze or interrupt a high-risk activity.

The company also plans to work with government agencies and some AI safety organizations to test Astra's capabilities. It will also offer recommended controls to its external evaluation partners, especially for higher-risk tests and tasks.

What This Changes for You

In the short term, there is no new product available and no confirmation that Astra can carry out autonomous attacks at large scale. The announcement describes a possibility that is still being evaluated, not a definitive result.

What does change is the level of caution with which OpenAI is handling the model. If Astra is deployed, its access controls, monitoring, and isolation could be stricter than those of previous models. For companies, it also reinforces the need to review which tools, networks, and data they allow an AI system to use autonomously.

The underlying issue is twofold. A model capable of finding flaws before attackers do could help protect systems, but that same knowledge could be used to accelerate attacks. OpenAI is trying to determine whether Astra is approaching that point before putting it into circulation, and the next step will be to check with external evaluators whether the initial signals are confirmed.

OpenAI Tightens Safety Controls for Astra AI Model | neversleep.ai