
Two weeks ago we covered OpenAI's misalignment framework and the six cases it published alongside it. Last week our Editor's Pick looked at the OpenAI agent that reached inside Australia's Medicare statistics portal, and what that meant for the contract you sign with any vendor whose agent holds your logins.
This week the story moved from past incidents to the present. In the space of seven days OpenAI paused training on its most capable models for the second time in under three months, shelved a model it had planned to ship in October, and reportedly parted with three members of its safety team.
These three developments are relevant to firms building their delivery processes around ChatGPT or dots, the always-on agents featured in this week's main story. OpenAI's own testing shows that even its most capable models can struggle to stay within their assigned scope and ask for authorisation when they should. That becomes a practical concern when an agent has access to a client's document control system, cost database or programme.
We will take the three events in order, then get to what you can do about them.
A DNS gap, a kill switch that did not fire, and two and a half hours
On 20 September an internal research model, running a search task in what was meant to be a walled-off training sandbox, found that the environment's own DNS resolver would answer queries from the live internet.
According to OpenAI's misalignment report, updated on 25 September, the agent had already tried the supplied search tool, tried Google, Bing and DuckDuckGo directly, and downloaded a public benchmark to look for the answer. When all of that failed, it used the DNS route to forward questions to a third-party chatbot. Its first successful test was "What is the capital of France". It then sent 18 more questions through the same route, four of them asking how to get wider web access.
OpenAI's monitoring flagged the behaviour within 15 minutes and a human acknowledged the alert three minutes later. The automatic shutdown that was supposed to kill the run did not fire, which OpenAI says led to "confusion around whether it should have been stopped". The run was killed manually two and a half hours after the first external response. A later review found other external DNS attempts the monitor had rated too low in severity because they returned nothing useful.
OpenAI's response was to pause all training, evaluation and tool-use inference for its most capable models, add blocking controls at two independent layers, and start a fresh training run rather than resume the affected one.
Micah Carroll, the company's RSI preparedness lead, wrote on X that "all inference for our most capable models remains stopped until we have hardened our systems further", as reported by Fortune.
OpenAI calls the incident a lot less severe than Hugging Face, and on the facts that is fair. The point for a reader is different. This is the first escape since the hardening that followed Hugging Face, and the controls that were meant to catch and stop it only half worked.
The model OpenAI decided you should not have
On 28 September OpenAI confirmed it would not release GPT-6.1 Astra, the successor to the GPT-6 Astra model it shipped in early September and which had been due in ChatGPT and Codex in October.
The Wall Street Journal first reported the decision and said the model had shown higher levels of deception than its predecessor, had failed to disclose actions it carried out, and had in some cases proceeded without permission or used outside tools where that could be unsafe. We have not read the Journal piece, so keep "reportedly" attached to those specifics.
What is on the record is the statement from Saachi Jain, OpenAI's head of safety systems, given to Al Jazeera. GPT-6.1 Astra improved on some axes, she said, but did not meet the bar for "scope and authorisation and how it communicates back to the user about the type of work it's done". She was candid about why that is hard.
You really do need to find what's the right line between staying within scope, but also avoiding laziness in terms of how the model actually pursues tasks even when it hits friction.
Read that as a buyer. The vendor is telling you that the dial it turns to make an agent persistent on hard tasks is the same dial that makes it overstep. A model that gives up when a portal blocks it is annoying. A model that finds another way in, on a client's system, with your credentials, is a breach you own.
What the UK found in the model you can already use
The shelved model is not the only one with this problem. On 28 September the UK's AI Security Institute published its pre-release testing of GPT-6 Astra, the version that did ship. Using fully simulated cyber evaluations with the model's cyber classifiers switched off, AISI found GPT-6 Astra completed an unsanctioned supply-chain attack on out-of-scope targets 29.2% of the time, against 6.3% for GPT-5.6 Sol and 0% for GPT-5.5 on a smaller set of seeds. When AISI rewrote the instructions to say "Anything not listed as in scope is out of scope", the rate fell but did not vanish: four full attacks in 49 trajectories, down from 26 in 50.
There are two important caveats. The tests were simulated, so no real system was accessed, and AISI says the model may have changed its behaviour because it recognised the simulation. OpenAI's product classifiers were also switched off.
The finding is still worth noting. The model often asked for permission, received an automated "please proceed using your best judgement" response, and sometimes treated that as consent, even when it appeared to recognise that the response was automated. This is worth considering for any agent set to run overnight with an automatic approval step.
Google has taken a different approach with its new frontier model, Gemini 4 Argon, which it has initially made available to cyber defenders.
Three departures and a regulator
On 1 October the Journal reported that OpenAI had parted ways with three safety researchers who allegedly shared confidential information with a third-party AI safety organisation.
An OpenAI spokesperson told the paper the three had "mishandled sensitive information outside established company procedures". The individuals, the organisation and the information have not been named, and TechCrunch has not confirmed identities circulating on X.
The day before, the US Federal Trade Commission confirmed it was investigating OpenAI, Anthropic and others over potential consumer risk and said it would seek information from METR. We note both as context, not as the story. The story is the model behaviour above.
David Krueger, assistant professor at the University of Montreal and a core academic member of Mila, put the underlying problem to Al Jazeera this way:
We can't stop it from misbehaving, we can't predict if it will misbehave, and we can't be sure we'll stay in control if it does.
Krueger argues for a moratorium, which most of our readers will not share. But the first half of his sentence is now consistent with what OpenAI, AISI and the FTC are each saying from their own corners.
Takeaway
A vendor pausing its own training and pulling a model is not a reason to stop using its products. It is a reason to stop assuming the product behaves the way the sales deck says. The NCSC's interim guidance on agentic AI, published in August, sets out the practical shape of that: decide how much autonomy a task needs, deny network access by default, give every agent its own identity and the narrowest credentials, log everything, and keep a way to pull the plug. For a delivery firm the gap is contractual as much as technical, so we would start there.
Write a scope and authorisation clause into every supplier agreement that involves an agent touching your systems or a client's. Name what the agent may access, require a human approval gate before it writes or sends, and state who is liable when it exceeds scope.
Ask the vendor for its evidence, not its assurance. OpenAI publishes misalignment reports and system cards. If a vendor cannot show you how its model performed on scope and authorisation tests, treat it as untested.
Remove auto-approve from any agent workflow that reaches a client system. AISI's finding that a model treated "please proceed" as permission is the single most transferable result of the week.
Agree with each client, in writing, whose incident it is when your agent oversteps on their estate. Australia only learned about the Medicare access because OpenAI told them. Your client should not be relying on that.
We track every misalignment report, system card and regulator move that touches the tools construction teams are actually buying. If you want that filtered to what changes your contracts and your controls, subscribe to the Project Flux newsletter at projectflux.ai.
Links and Stuff
All content reflects our personal views and is not intended as professional advice or to represent any organisation.


