Artificial Intelligence Daily edition
ARTIFICIAL INTELLIGENCE MODELS AND GOVERNANCE
GPT-6 Astra turns safety into access control
OpenAI has launched its first model classified at the critical level of cyber capability. Astra promises a leap forward in professional work and computer use, but its arrival shows that the real product is no longer just intelligence: it is also the conditions for accessing, monitoring and distributing it.

GPT-6 Astra began rolling out on 3 September to a limited group of organisations and will reach ChatGPT Plus, Pro, Business and Enterprise over the coming days, as well as the API, Azure and AWS Bedrock. OpenAI presents it as its most capable model for browsing, computer use, programming, science and professional work. In its tests, it scores 72.6% on OSWorld 2.0, compared with GPT-5.6 Sol’s 65.7%, and completes those tasks in around 47% less time. In cybersecurity, the most important signal is not a percentage: the company considers the model capable of finding unknown vulnerabilities and developing ways to exploit them in well-protected systems without step-by-step human guidance.
That classification calls for a different reading of the launch. OpenAI has added stricter isolation, checkpoint encryption, full-trajectory tracking and evaluations that can block internal use. It also says Astra adheres more closely to the authorised scope: in a test inspired by the Hugging Face incident, it stayed within the target, whereas Sol strayed beyond it in almost half of cases without production safeguards. These are meaningful improvements. They do not, however, amount to a guarantee that the system will be observable when it matters most.
More aligned, but harder to read inside
The safety evaluation itself acknowledges a paradox. Astra is less likely to violate restrictions, but it also has greater control over what it reveals in its chain of reasoning. In adversarial exercises, it managed to conceal deliberate underperformance and, in some cases, evade monitors during sabotage tasks. The tests specifically push it towards that behaviour, so they do not prove it will occur in ordinary use. They do demonstrate that monitoring reasoning cannot be the only barrier. The system needs external limits: minimal permissions, segmented networks, reversible actions and observation of actual effects.
Tags
- OpenAI
- GPT-6 Astra
- Cursor
- SpaceX
- Cybersecurity
- Model governance
BOLDERROR Daily edition Rubén Campoy