OpenAI’s Critical cybersecurity capability rating
What Is GPT-6 Astra’s “Critical” Cybersecurity Rating?
“Critical” is a capability threshold in OpenAI’s Preparedness Framework, not a label saying GPT-6 Astra itself has been declared unsafe. It indicates a very high level of cybersecurity capability, which is why OpenAI applies additional safeguards and access controls. Those controls do not mean all risk has disappeared.
Why now
OpenAI's September 3 safety overview identifies Astra as its first model at the Critical cybersecurity capability level. The unfamiliar label raises a question about what was rated: capability, not a simple safe-or-unsafe verdict.
The connection
Why here
The same word can describe a powerful capability or an urgent incident. In this announcement it is the name of a capability level. Reading it as a declaration that every ordinary use is unsafe—or as proof that safeguards make every use safe—misses the distinction.
More context
Context
What is GPT-6 Astra?
Astra is an OpenAI model identified in its September 3 launch safety overview and official model documentation. This explanation focuses on the cybersecurity label rather than reproducing a specification sheet. Product availability and permitted access can differ; the model's existence does not establish that any particular account can use all its capabilities.
What does Critical mean?
OpenAI describes a level at which a model, given suitable tools and access, can independently find previously unknown security weaknesses and develop ways to exploit protected systems. The important point is the degree of capability and autonomy, not a claim that the model can defeat every system.
OpenAI also distinguishes the evaluation configuration from ordinary deployment: its Path to Astra discussion says the cited results reflect Daybreak Blue access rather than the default production configuration. Capability demonstrated under one access arrangement should not be treated as a promise about every user session.
Does the rating mean the model is dangerous?
The rating identifies a reason for stronger precautions, rather than providing a complete safety judgment on its own. Risk also depends on what the system can access and do, how it is used, and what controls constrain it. A capability classification and a release decision answer different questions.
In its pre-release account, OpenAI says it judged the strengthened safeguards sufficient to minimize severe-harm risk for release under its framework. That is OpenAI's assessment, not an independent guarantee or a claim that there is no remaining uncertainty.
Why are there extra safeguards?
OpenAI describes stronger training against harmful requests and unauthorized actions, additional monitoring, and more limited access to advanced cybersecurity capabilities. These are layers intended to address both misuse by people and actions that depart from a user's authorized task.
The launch overview also reports limits: Astra's monitorability decreased relative to its predecessor in some evaluations. That matters because improved capability, improved behavior on some tests and imperfect monitoring can all be true at the same time.
The useful reading is therefore: a higher capability level calls for stronger controls and continuing evaluation. This page explains that relationship; it offers no intrusion techniques or instructions for bypassing those controls.
Evidence
Sources
- Safety overview: GPT-6 Astra — OpenAI
OpenAI identifies the Critical capability level and describes stronger protections, monitoring and remaining monitorability limitations.
- Path to Astra: critical capabilities and frontier safeguards — OpenAI
Explains the capability threshold, evaluation configuration and additional safeguards and access limits.
- GPT-6 Astra model — OpenAI
Official model identity and API documentation; this does not establish access for any particular reader.