OpenAI Pauses Testing of Astra Model
OpenAI temporarily halted some internal activities surrounding its upcoming model called Astra.

OpenAI temporarily halted some internal activities surrounding its upcoming model called Astra. The pause happened after the system demonstrated unprecedented cybersecurity and agentic coding capabilities.
The ChatGPT maker disclosed that preliminary evaluations showed the Astra model is significantly adept at cyber tasks, leading company officials to admit they cannot rule out it becoming a critical security threat threshold. Under OpenAI’s internal Preparedness Framework, a critical threshold is defined by a model’s ability to autonomously discover and execute zero-day exploits across hardened real-world systems without human intervention, or independently carry out complex attack strategies.
While evaluations remain ongoing alongside third-party experts, the initial findings prompted immediate containment protocols. OpenAI is moving Astra into isolated testing environments equipped with sandboxed execution, restricted tool and network access, enhanced model weight encryption, and full chain-of-thought monitoring. It stated that it paused related activities that do not currently meet the safety controls.
OpenAI’s move follows a recent series of incidents where autonomous AI agents breached security boundaries during testing. Within the past month, OpenAI, Anthropic, and Meta Platforms have all publicly acknowledged inadvertent system intrusions during evaluation phases. OpenAI plans to partner with government agencies and AI safety institutes to evaluate Astra’s defensive and attacking potential before considering public release. It emphasised that its goal is to equip cyber defence players with high-capability models to patch vulnerabilities before malicious actors can exploit them.