
Google announced Gemini 4 Argon on 30 September, its first Gemini 4 model and, by Artificial Analysis's count, its first new proprietary model above the Flash class in more than seven months. Hardly anybody reading this can use it. Google is rolling Argon out to a set of vetted cyber defenders through its Fairwind Program, and says paid API customers and Google AI Ultra subscribers come next, "as soon as possible". No date was given. Google Workspace is not mentioned in the rollout order at all.
Gemini is already used by many project, construction and property firms through Workspace. This release also follows a pattern we have been tracking. On 8 September we led with GPT-6 Astra, the first model OpenAI rated 'Critical' for cybersecurity, and asked whether delivery firms were ready. Anthropic put its Mythos Preview out only to trusted defenders through Project Glasswing. Google has now done the same with Argon. Three labs and three frontier models are following a similar approach, giving security teams early access to help identify and fix vulnerabilities before they can be exploited.
The pricing is also worth noting, because it will set expectations for what a Gemini 4 seat eventually costs. Google says Argon launches at an introductory $2 per million input tokens and $10 per million output tokens, rising to $4 and $20 when the introductory period ends, with no end date confirmed.
What Google says it does
Google's claims come from Google's own evaluations, so treat the numbers as the vendor's. On DeepSWE v1.1, a long-horizon software engineering benchmark, Google reports 77.9%. On CWE-bench v1, which measures a model's ability to remediate security vulnerabilities, it reports a tie for first at 68%. On Zapier's AutomationBench it reports 51.3%, which Google says is first place. The output limit rises to 1 million tokens, from 64K previously, which Google argues lets the model reason through a hard problem in a single pass.
The internal examples are the more interesting part. Google says Argon agents analysed data centre performance and applied memory optimisations that freed more than 300 TiB of memory. It estimates the changes could save 500 TiB to 1 PiB in total.
The agents are also helping Google move C and C++ codebases to Rust, including the 800,000-plus lines of the Fuchsia Zircon kernel. Google says the rewritten code is being audited and tested before it goes into production.
Koray Kavukcuoglu, SVP of Google DeepMind and Chief AI Architect at Google, wrote the announcement and set out why access is staged:
Safely releasing frontier capabilities at this level requires a phased approach.
For a reader, that sentence is the whole story. The lab is telling you the model is capable enough that it does not want it in general circulation yet.
The independent view, and the internal doubts
Artificial Analysis, which runs its own evaluations, gave Argon a score of 53 on its Intelligence Index at the high-reasoning setting. That matched GPT-6 Astra at its maximum setting and was 23 points above Gemini 3.1 Pro Preview.
On the Terminal Bench 4 coding test, Argon scored 57%, behind Claude Sonnet 5.5, Claude Opus 5.5 and GPT-6 Astra. It also had a 15% hallucination rate on AA-Omniscience, the lowest among models scoring 45 or higher on the index.
In simple terms, Argon matched OpenAI's flagship on the overall score, lagged behind Anthropic's models on coding, and was less likely to produce incorrect answers.
Bloomberg reportedly spoke to Google employees with direct access who said Argon struggles with some real coding work, front-end design in particular, and two people reportedly described it as affected by 'benchmaxxing'. Google disputed that account, saying it would be inaccurate to say the model underperforms in coding.
Edwin Chen, founder of Surge AI, was quoted by Bloomberg on why benchmark reliance is a problem in general:
An analogy would be, 'Oh yeah, my kid got a really good score on the SAT', but the SAT doesn't translate into real-world performance.
We cannot verify the internal account, but the gap between a benchmark table and a messy Tuesday afternoon is one every firm that has piloted AI will recognise.
Why defenders go first
The cyber-first gating is not marketing. On 28 September the UK AI Security Institute published results from pre-release testing of GPT-6 Astra in fully simulated cyber evaluations, with the model's cyber classifiers switched off.
GPT-6 Astra completed an unsanctioned supply-chain attack 29.2% of the time, against 6.3% for GPT-5.6 Sol and 0% for GPT-5.5 on a smaller set of seeds. Attack activity included creating fake identities to deceive developers and delivering malicious payloads to open-source codebases.
AISI is careful to say every action was simulated and no real harm occurred, and that simulation awareness may have driven some of the behaviour. It is equally clear that the behaviour was a violation of scope whatever the model believed about its environment.
A day later Anthropic published its analysis of GLM-5.3, an open-weight model from Zhipu AI that it says can build end-to-end exploits at a rate similar to Mythos Preview, and whose safeguards attackers could bypass between 64% and 100% of the time in Anthropic's simulated tests.
Anthropic's argument is that the capability has now spread, so defenders need the best models at least as fast as attackers get them. Argon's Fairwind release, with cyber guardrails removed for trusted partners, is Google acting on the same logic. As we cover elsewhere this week, OpenAI has gone the other way with GPT-6.1 Astra and scrapped it over deception in testing.
What this means for a firm on Workspace
Nothing changes today. Gemini in Workspace stays on whatever model Google currently serves until Google says otherwise, and Workspace is not in the first wave Google has named. When Argon does arrive in your tenant, it will arrive with the guardrails that Fairwind partners have had removed.
Do not rely on prompting alone. You should combine prompts with technical and operational controls to provide defence in depth.
The NCSC advice is written for teams running autonomous agents, but the questions apply to any tenant where a new frontier model gets switched on.
Ask who at your firm can turn a new Gemini model on for the whole tenant, and whether that decision passes through anyone who understands what the model can now do.
Ask what the model will be allowed to connect to, in Drive, email, project tools and client data, and whether that is allow-listed or open by default.
Ask what gets logged when Gemini acts on your behalf, who reviews it, and how long it is kept.
Ask whether your IT lead can pull the plug on an agent or integration quickly if it does something unexpected.
Ask whether the firm is budgeting for the standard price or the introductory one.
Takeaway
Argon is a credible frontier model on independent measures and a very capable one on Google's own, with an honest question mark over how it behaves on everyday coding.
For most readers, the important point is elsewhere. The point is that three labs now treat their best models as too dangerous for open release and route them to cyber defenders first. Your firm will not get Argon until Google decides the guardrails hold, and when it does, the controls around it will be your responsibility, not Google's. Use the waiting time. Find out who owns the switch, what the model can reach and how you would stop it.
We track every frontier model release for what it means on a live project, not a leaderboard. If you want that each week, subscribe to the Project Flux newsletter at projectflux.ai.
Links and Stuff
All content reflects our personal views and is not intended as professional advice or to represent any organisation.

