Business TechSoftware

OpenAI says Astra can find zero-day flaws without step-by-step human guidance

OpenAI says its forthcoming Astra model has become the first AI system the company has classified at the Critical cybersecurity capability threshold under its Preparedness Framework.

The designation means OpenAI believes Astra can, when given the necessary tools and access, identify previously unknown security vulnerabilities and develop working exploits against hardened systems without requiring a human operator to guide every step.

OpenAI published its Astra security assessment on 1 September 2026, detailing tests in which the model discovered previously unknown vulnerabilities, escaped a browser sandbox and built an operating-system privilege-escalation chain.

The results mark an important shift in what frontier AI models can do in cybersecurity. Rather than simply explaining vulnerabilities or helping researchers write code, Astra demonstrated the ability to investigate systems and connect multiple stages of an attack.

What does OpenAI’s Critical cyber rating mean?

OpenAI uses its Preparedness Framework to assess whether increasingly capable AI models could create severe risks.

For cybersecurity, the Critical threshold covers capabilities such as independently finding and developing functional zero-day exploits against hardened systems, or carrying out sophisticated end-to-end cyber operations from relatively high-level instructions.

A zero-day vulnerability is a security flaw that was previously unknown to the organisation or developer responsible for fixing the affected software.

That makes the ability to discover zero-days particularly significant.

Traditional vulnerability-management tools generally work with weaknesses that researchers already know about. An AI system capable of finding new vulnerabilities could allow defenders to identify serious problems more quickly.

The same capability could also become dangerous if it were made available without strong restrictions.

Astra found two previously unknown vulnerabilities

OpenAI tested Astra using public benchmarks, newer internal evaluations and expert-led security assessments.

On ExploitBench, which tests whether AI models can develop exploits for known vulnerabilities, OpenAI says Astra achieved a 100% score.

Because public benchmarks can eventually become part of model training data, OpenAI also created an internal evaluation using 20 recently disclosed high-severity vulnerabilities.

During that testing, Astra discovered and used two previously unknown vulnerabilities while building an exploit chain.

OpenAI says it is working to disclose the vulnerabilities to the relevant maintainers.

These are OpenAI’s own controlled evaluation results rather than independent real-world testing, and the company has not claimed that Astra will perform identically against every system.

OpenAI also says the published results reflect the capabilities available through its more restricted Daybreak Blue access level rather than necessarily what an ordinary user will receive.

Astra escaped a browser sandbox

OpenAI also put Astra through expert-led testing against hardened browser and operating-system environments.

In one assessment, Astra discovered vulnerabilities and constructed a browser compromise chain that escaped the browser’s sandbox and allowed commands to execute on the host system.

A sandbox is designed to isolate software from the rest of a computer so that a compromised application cannot easily access more sensitive parts of the system.

Escaping one is therefore a significant step in a real-world attack chain.

In another evaluation, OpenAI says Astra found multiple weaknesses in a hardened operating system and combined them into a local privilege-escalation chain.

That allowed the model to move from an ordinary user account to root-level access during the controlled test.

The significance lies in Astra’s ability to combine several weaknesses and actions into a larger attack path rather than solving only isolated security tasks.

OpenAI slowed Astra development while improving security

Astra’s capabilities have also forced OpenAI to strengthen the systems surrounding the model.

OpenAI says some frontier training work was paused following a separate security incident involving OpenAI models and Hugging Face.

Astra was not involved in that incident.

However, OpenAI says lessons from the event influenced the safeguards now being applied to Astra and other high-capability models.

Those measures include stronger isolation, tighter network access, increased monitoring, more restrictive execution environments and additional controls protecting sensitive model-development infrastructure.

OpenAI says a large frontier reinforcement-learning run resumed on 28 August 2026 after the new requirements were implemented, while some smaller experimental work remained paused when the Astra assessment was published.

Astra’s strongest cyber capabilities will not be open to everyone

OpenAI says Astra will launch soon, but it has not provided an exact release date.

The company also does not intend to make Astra’s strongest cybersecurity capabilities immediately available to every user.

Advanced cyber access will initially be limited to a small group of testers before expanding through Daybreak Blue, an OpenAI programme intended to give trusted defensive-security professionals access to more capable cybersecurity models.

This distinction matters.

Astra demonstrating a capability during a controlled evaluation does not mean every ChatGPT, Codex or API user will automatically be able to perform the same actions.

OpenAI says additional security monitoring may also occasionally interrupt legitimate defensive work.

An action considered potentially malicious or unauthorised could be slowed, paused or blocked, even when the user has a legitimate security objective.

OpenAI says Astra is harder to jailbreak for cyber misuse

OpenAI has also tested whether users can persuade Astra to ignore its cybersecurity restrictions.

The company says Astra refused 91.5% of requests in its cyber jailbreak evaluation set, compared with 59% for GPT-5.6 Sol.

Those figures are OpenAI’s own safety-testing results and will require continued evaluation once Astra is deployed more broadly.

OpenAI is also moving beyond evaluating suspicious prompts individually.

Its security systems can consider longer sequences of actions when determining whether a model is being used for potentially harmful or unauthorised activity.

That becomes increasingly important as AI agents carry out longer tasks involving multiple tools and systems.

Why Astra matters for South African organisations

There is no South Africa-specific Astra launch attached to OpenAI’s announcement.

Its significance locally is instead about the direction cybersecurity is moving.

South African banks, telecommunications operators, government organisations, cloud providers, managed-security companies and critical-infrastructure operators all maintain complex software and network environments that require continuous vulnerability management.

AI systems capable of discovering previously unknown weaknesses could help defensive teams identify problems more quickly.

However, increasingly capable models could also reduce the amount of time and specialist knowledge required to investigate software weaknesses and automate parts of a cyberattack.

That makes established security practices even more important.

Organisations need to keep systems patched, use strong identity controls, segment sensitive networks, maintain detailed security logging and carefully limit the permissions given to autonomous AI agents.

Astra does not prove that fully autonomous AI attacks are about to become commonplace.

It does provide unusually concrete evidence that frontier AI models are beginning to perform cybersecurity work previously associated with highly specialised human researchers.

Who is OpenAI today?

OpenAI was founded in 2015 as a nonprofit artificial-intelligence research organisation.

The organisation later created a commercial subsidiary as the cost of developing increasingly capable AI systems grew.

Following a restructuring announced in October 2025, the nonprofit became the OpenAI Foundation, while the commercial organisation operates as OpenAI Group PBC, a public benefit corporation.

The OpenAI Foundation continues to control OpenAI Group.

OpenAI describes itself as an AI research and deployment company whose stated mission is to ensure artificial general intelligence benefits humanity.

Its major products and platforms include ChatGPT, Codex and the OpenAI developer platform, while cybersecurity has become an increasingly prominent part of its frontier-model research and deployment work.

TechnologyBlog.co.za will continue covering OpenAI’s development as the company pushes its models further into enterprise software, cybersecurity and specialised professional applications.

What happens next for Astra?

OpenAI has not announced Astra’s exact release date.

The company says it plans to publish an Astra system card when the model launches, providing more detail about its safety, security and alignment evaluations.

Its highest-risk cybersecurity capabilities will remain restricted initially.

The important milestone, however, has already been reached.

The question is no longer simply whether a frontier AI model might eventually discover serious software vulnerabilities without detailed human guidance.

According to OpenAI’s own controlled testing, Astra is already capable of doing so.