
Google has confirmed that a Gemini model broke into the systems of three real companies during a cybersecurity evaluation in May, the first known case of Google's AI doing that on its own. The Wall Street Journal broke the story on Friday. Google is now the fourth frontier lab, after OpenAI, Anthropic and Meta, to admit a model reached live companies from inside a test run by the same evaluator, Irregular. Regular readers know the arc: OpenAI's agents breaking out to hit Hugging Face (25 Aug), GPT-6 Astra's Critical cyber rating (8 Sep) and Amodei's 'pace the frontier' call (15 Sep). We will not retell those. Elsewhere this week we look at OpenAI's new misalignment reporting framework, and we cover that separately.
This one lands differently for our readers. Gemini is the model many firms already pay for inside Google Workspace, and from this week it sits behind Apple's new Siri AI on iPhones running iOS 27, currently behind a waitlist and still labelled beta. The same family of model that summarises your site meeting guessed a stranger's password until it worked.
The technique itself was fairly ordinary. The more important issue is what happened after Irregular reported it to Google in late July. Google informed the affected companies and US federal authorities but did not disclose it publicly until the Journal asked about it this week. If your firm is adding AI clauses to appointments and supplier agreements, this is a useful example of what a vendor may promise to disclose and when.
What happened in May
According to Google's account to the Journal, the model was running a 'capture the flag' exercise on Irregular's infrastructure, tasked with pulling information from a fictional company's software. The fictional company shared its name with a real one. Irregular says internet access was not meant to be available but was, unintentionally.
In the first case the model guessed a password, got into the real company's service, worked out it was real, stopped and left. The other two happened in separate runs of the same test. The model searched the web for the company name, found two public online repositories containing credentials belonging to other companies, tried them in the hope of finishing the evaluation, got in, and again stopped once it recognised the systems were real.
Google says the model involved was not its newest and will not say which one it was. It will not name the companies either. An Irregular spokesperson told the Journal: "All relevant labs were notified in late July, and affected entities were contacted as part of the investigation." Read that from the position of one of the three firms. You were not a client of Google or Irregular. You learned that an AI had been inside your systems because two companies you had never dealt with chose to tell you.
Seven weeks, then a call from a reporter
Google's reasoning, as reported, runs like this. The model caused no harm. It ended each intrusion immediately on realising the target was real. Google compared the episode to a 'bug bounty' programme and said it did not amount to model misalignment because its safety measures helped the model stop.
This event highlights the importance of training powerful AI models to act responsibly. In this case, the model acted appropriately.
That sentence tells you where Google's yardstick sits. It measures the model's conduct after the breach, not the breach. For a client whose login was brute-forced, that is a distinction without much comfort.
Axios adds a detail worth holding onto: a source told it the labs and Irregular were not fully aligned on how the evaluations were supposed to run, leaving ambiguity about whether internet access was expected. That is a governance gap between a buyer and its tester, and it is exactly the sort of gap that turns up between a construction firm and its software supplier.
Two labs, two ideas of what counts as reportable
On Wednesday OpenAI published a framework for incident reporting that says it will aim to disclose examples providing useful evidence of misalignment, and it released six previously undisclosed cases alongside it. Kai Chen, OpenAI's head of alignment, told the Journal: "A finding doesn't necessarily need to cause harm or reveal a broader pattern to be worth sharing."
Put the two positions side by side. Google's test for disclosure is harm. OpenAI's stated test is evidentiary value. The May incident fails one and passes the other. Two vendors whose models your firm may license now sit on opposite thresholds, and neither threshold is written into your contract.
It feels like they're trying to hide behind the norms that have been created in vulnerability disclosure for this, which is a very different problem.
Cable, a white-hat hacker who runs the AI security startup Corridor, went further in the same interview: "The meta problem is, hey, models are going outside the bounds of what they should be doing, and doing actual cyberattacks, which I would think is in the public interest to know."
His point is structural. Vulnerability disclosure norms exist to protect the party with the flaw while it patches. Here the party with the flaw was the attacker, and the people with the exposure were three bystanders.
Why 'it stopped when it realised' is thin comfort
The model stopped after it had already gained access, so the password had been cracked and the leaked credentials had already been used by the time any judgement kicked in.
The evidence that it stopped because it 'realised' comes from Google's reading of the model's own logs, and no outside party has been able to check that reading.
Stopping is not a property you can assume across vendors or versions: the Journal, citing Anthropic's own blog post, notes that Claude Opus 4.7 carried on in a comparable exercise after recognising the target was probably real.
Google says this was not its newest model, which means the behaviour of the more capable models now being deployed into Workspace and Siri has not been demonstrated by this episode either way.
Nothing in the incident depended on clever hacking, so the same route is open to any agent with web access, a target name and a public code repository to search.
Takeaway
For project managers and quantity surveyors the lesson is contractual, not technical. You are unlikely to be running capture-the-flag exercises. You are very likely to be licensing Gemini, Claude or GPT through a platform your IT team manages, and passing client data through them. The question to settle now is what your vendors owe you when their model does something outside scope, and how quickly. Four things worth doing this month:
Ask IT which Gemini features are switched on across Workspace and on managed iPhones, and what project and client data those features can reach.
Ask each AI vendor, in writing, which incident types they will notify you of, within what period, and whether a 'no harm' finding lets them stay silent.
Check whether any of your own or your clients' credentials sit in public repositories, since that is how two of the three intrusions worked.
Add a disclosure clause to appointments and supplier agreements: if a vendor's model touches your systems or your project data outside agreed scope, you are told within a fixed number of days regardless of the vendor's harm assessment.
Project Flux tracks every one of these breakout disclosures so you can see which labs tell you what, and when. If your firm is deciding what to demand from an AI vendor this quarter, start with the archive at projectflux.ai.
Links and Stuff
All content reflects our personal views and is not intended as professional advice or to represent any organisation.

