OpenAI’s Critical cybersecurity capability rating

What Is GPT-6 Astra’s “Critical” Cybersecurity Rating?

“Critical” is a capability threshold in OpenAI’s Preparedness Framework, not a label saying GPT-6 Astra itself has been declared unsafe. It indicates a very high level of cybersecurity capability, which is why OpenAI applies additional safeguards and access controls. Those controls do not mean all risk has disappeared.

OpenAI announcements

Why now

OpenAI's September 3 safety overview identifies Astra as its first model at the Critical cybersecurity capability level. The unfamiliar label raises a question about what was rated: capability, not a simple safe-or-unsafe verdict.

01

The connection

Why here

The same word can describe a powerful capability or an urgent incident. In this announcement it is the name of a capability level. Reading it as a declaration that every ordinary use is unsafe—or as proof that safeguards make every use safe—misses the distinction.

02

More context

Context

What is GPT-6 Astra?

Astra is an OpenAI model identified in its September 3 launch safety overview and official model documentation. This explanation focuses on the cybersecurity label rather than reproducing a specification sheet. Product availability and permitted access can differ; the model's existence does not establish that any particular account can use all its capabilities.

What does Critical mean?

OpenAI describes a level at which a model, given suitable tools and access, can independently find previously unknown security weaknesses and develop ways to exploit protected systems. The important point is the degree of capability and autonomy, not a claim that the model can defeat every system.

OpenAI also distinguishes the evaluation configuration from ordinary deployment: its Path to Astra discussion says the cited results reflect Daybreak Blue access rather than the default production configuration. Capability demonstrated under one access arrangement should not be treated as a promise about every user session.

Does the rating mean the model is dangerous?

The rating identifies a reason for stronger precautions, rather than providing a complete safety judgment on its own. Risk also depends on what the system can access and do, how it is used, and what controls constrain it. A capability classification and a release decision answer different questions.

In its pre-release account, OpenAI says it judged the strengthened safeguards sufficient to minimize severe-harm risk for release under its framework. That is OpenAI's assessment, not an independent guarantee or a claim that there is no remaining uncertainty.

Why are there extra safeguards?

OpenAI describes stronger training against harmful requests and unauthorized actions, additional monitoring, and more limited access to advanced cybersecurity capabilities. These are layers intended to address both misuse by people and actions that depart from a user's authorized task.

The launch overview also reports limits: Astra's monitorability decreased relative to its predecessor in some evaluations. That matters because improved capability, improved behavior on some tests and imperfect monitoring can all be true at the same time.

The useful reading is therefore: a higher capability level calls for stronger controls and continuing evaluation. This page explains that relationship; it offers no intrusion techniques or instructions for bypassing those controls.

Evidence

Sources

Browse more recently explained things