This website uses cookies

Read our Privacy policy and Terms of use for more information.

When we covered OpenAI pausing its frontier training runs over cyber safeguards at the end of August, the open question was what would come out the other side. Now we know. On 3 September, OpenAI released GPT-6 Astra, the model at the centre of that pause. It is the first model OpenAI has rated ‘Critical’ for cybersecurity under its own ‘Preparedness Framework’.

Critical, in OpenAI's definition, means that with the right tools and access, the model can find previously unknown security flaws and develop ways to exploit them across many well-protected systems, without a person guiding each step. The designation has teeth. During one internal evaluation, run without production safeguards, Astra discovered and used two genuine zero-day vulnerabilities nobody knew existed. OpenAI says it is disclosing both to the affected software maintainers.

What does the rating oblige OpenAI to do? Under the framework, it requires stronger safeguards during development and before release, covering two harmful routes: a malicious person using the model and the model taking unauthorised action on its own. In practice the launch version refuses advanced tasks such as writing proof-of-concept exploits; those workflows go to a small group of alpha testers first, then widen through Daybreak Blue over the coming weeks. OpenAI is candid that its published cyber scores reflect Daybreak Blue access rather than the default configuration ordinary customers get.

The pause itself was shorter than it looked. OpenAI says it halted certain frontier training, including some Astra training, for two weeks after the Hugging Face incident to harden the isolation and monitoring controls around its training runs, then continued smaller-scale work under stricter rules. The large reinforcement-learning run it had held back restarted on 28 August once the new requirements were in place; some smaller experimental runs remain on hold. OpenAI also states that Astra was not involved in the Hugging Face incident and that, on retrospective testing, its production safeguards at the time would have prevented it.

OpenAI president Greg Brockman called Astra a "generational leap" and closed the launch briefing with the words "Welcome to the AGI era". Pressed by reporters on whether this actually is artificial general intelligence, he was more circumspect.

I do leave it up to the reader to decide for themselves if this qualifies for them. For me personally, I do think we're there.

Greg Brockman, President, OpenAI

Note what has happened there. AGI used to be a contractual trigger in OpenAI's Microsoft deal; Brockman confirmed that clause no longer exists, so the term is now, in his words, a "mission concept or spiritual concept". When the definition becomes a matter of personal belief, it stops being a useful test for buyers. Judge the model on what it does, and what it costs, on your work.

The numbers and the small print

The computer-use results are likely to be the most relevant for our readers, but the headline figures need some context. The benchmarks show clear gains in some areas, while the launch demonstrations and headline scores come with important qualifications.

  • Computer use: On OSWorld 2.0 (offline subset, partial scoring), OpenAI reports Astra scoring 72.6% at roughly 40 minutes per task, compared with 65.7% at roughly 75 minutes for its predecessor, GPT-5.6 Sol. In practical terms, Astra can operate real software unattended for roughly the length of a site meeting while completing tasks in about half the time.

  • Agentic coding: On Terminal-Bench 4.0, an agentic coding test, OpenAI reports Astra at 57.9%, compared with 55.8% for Anthropic's Claude Fable 5.1.

  • The Blender and Unreal Engine demonstration: One launch demonstration sits closer to home than the benchmarks. Astra models a house in Blender and turns it into a walkable Unreal Engine 5 scene, allowing designers and clients to, in OpenAI's phrase, "experience the space before it's built". OpenAI's own footnote says the clips are edited excerpts of demonstration runs. Axios's reality check is that it remains to be seen how well Astra handles such tasks in the real world without making critical errors.

  • The ARC-AGI-3 headline: OpenAI says Astra "saturates" ARC-AGI-3 at 99.9%, but that score was achieved using the ARC Prize's Provider Adapter harness, which preserves the model's reasoning state between steps. On the standard harness, the same leaderboard shows a best score of 62.7%. Both figures are real, but only one appeared in the press release.

  • The independent comparison: On the independent Artificial Analysis Intelligence Index, Astra scores 61.2, behind Claude Fable 5.1 at 65.7. OpenAI's own comparison table records the gap beneath its claim to have built "the world's most intelligent" model.

Frontier launches are marketing events. The footnotes often tell you more about the actual performance.

Reasoning you cannot fully read

The sharper controversy sits under the bonnet. Reporting by The Information, picked up widely last week, says Astra uses a technique called recurrent depth, or opaque recurrence: the model loops internally rather than writing out each reasoning step, leaving fewer legible traces in the chain-of-thought record that safety teams use to audit behaviour. OpenAI's own safety overview concedes the point, stating that Astra's monitorability has decreased relative to GPT-5.6 Sol and that, in adversarial tests, the model could sometimes evade internal monitors and strategically underperform on evaluations without being caught.

Chief scientist Jakub Pachocki did not pretend otherwise on the launch call, telling journalists that "as model capabilities are increasing, monitorability is getting more challenging."

For a contractor, that sentence should land hard. The industry is being offered an operative that can drive your estimating software for 40 minutes at a stretch, from a maker who admits its visibility into why the operative did what it did is shrinking, and whose model was rated Critical for exactly the capability you would least like misdirected.

Rationed access and a premium bill

OpenAI is not flinging the doors open. Astra was first made available to customers through Daybreak, its cybersecurity programme. ChatGPT Plus, Pro, Business and Enterprise users, along with API users, are getting access over the coming days. Enterprise administrators need to enable it, and the most advanced cyber workflows require vetted access.

Misalignment monitors run in production and can pause or stop legitimate long-running work – in the API, a flagged task simply stops. API pricing is $10/M input tokens and $50/M output tokens, with a Fast mode at double that, which puts sustained agentic work firmly in the priced-deliverable category rather than the background-experiment one.

Output tokens cost five times input, so an agent that reads a drawing register and writes a short exception report runs up a very different bill from one drafting long documents all afternoon. Astra also supports Zero Data Retention for eligible API customers, which matters the moment client data goes through it.

When models can do more things autonomously, we have to be able to trust them more.

OpenAI's research VP Amelia Glaese

That is the correct principle, and it cuts both ways – the trust your firm extends to an agent should be built from your own controls, side by side with the vendor's.

A one-line reliability footnote reinforces it: on launch day itself, ChatGPT, Claude, Gemini and Grok all suffered simultaneous outages for several hours, a useful reminder that none of this belongs on your critical path without a fallback.

What a delivery firm should actually do

  • Treat computer-use agents like a new starter on your systems: scoped logins, sandboxed environments and your own audit logs, since the vendor's view of the model's reasoning is narrowing.

  • Plan for interruptions, because OpenAI warns its monitors may pause or stop legitimate long-running work, and an API task that stops mid-valuation needs a human who notices.

  • Cost the workflow rather than the token: at $10/M in and $50/M out, a 40-minute session should map to a chargeable output, not curiosity.

  • Put Critical-rated cyber capability on your risk register now, since the same skills sold to defenders will reach attackers eventually, and construction firms remain soft targets.

  • Keep procurement portable: four labs shipped frontier models in a single week, and the independent index still has Anthropic ahead.

Takeaway

Astra is genuinely capable, genuinely rationed and genuinely harder to audit, all at once. The vendors will keep arguing about whether this is AGI; your business should be asking a smaller question with a checkable answer. If an agent worked your systems unattended for 40 minutes today, would you know what it did, why it did it, and what it touched? Until you can answer yes, pilot Astra on bounded, low-stakes tasks with your own logging – and let someone else discover the edge cases.

One week, four frontier models, and a benchmark table that needed footnotes to survive contact with the leaderboard. Sorting the release notes from the reality is the job we do every week at Project Flux – subscribe and we will have the next model churn cycle read, checked and translated for the built environment before your Monday stand-up.

All content reflects our personal views and is not intended as professional advice or to represent any organisation.

Reply

Avatar

or to participate