Google opens Gemini 4 Argon to Fairwind defenders
Google published Gemini 4 Argon: our next era of frontier intelligence on 30 September 2026 US time (1 October in Singapore). The post is signed by Koray Kavukcuoglu, SVP of Google DeepMind and Chief AI Architect at Google.
Google says Argon "is rolling out to a set of trusted cyber defenders" through its Fairwind Program. Developers, enterprises and consumers come later, with no date given. Google says the wider release will start "with paid API customers and Google AI Ultra subscribers."
The post sets an introductory price of $2 per million input tokens and $10 per million output tokens. Cached input is priced at 95% off the input rate, which works out to $0.10 per million tokens by our arithmetic. A footnote says $4 input and $20 output per million tokens apply after the introductory period. The post does not say how long that period lasts.
Fairwind partners get Argon without cyber guardrails
Google says it will release Argon "without cyber guardrails" to trusted defenders and its own internal teams. The Fairwind Program page says Google works with over 650 partners and that a set of them get exclusive access to Argon.
The program page lists these access terms:
- Partners may carry out only dual-use tasks such as authorized threat simulation, reverse engineering and malware analysis, for defensive and academic research.
- Partners must use user-level authentication, phishing-resistant MFA and access controls.
- Partners may give Argon access only to internal cybersecurity, incident response or penetration testing teams, and must track employee use.
- Partners may not share, redistribute or sell access.
- Google runs background checks on applying organizations.
The page says Argon supports zero data retention when used as a managed model on Gemini Enterprise. It says Google gives priority to governments, critical infrastructure operators and core technology platforms.
Google's own benchmark table for Gemini 4 Argon and three rival models.
Source: Google, Gemini 4 Argon: our next era of frontier intelligence, 30 September 2026
A 1M output limit and Google's benchmark table
Google says Argon's output token limit rises to 1M tokens, up from 64K. The scores below come from Google's post and its evaluation methodology PDF. The post cites no independent reproduction of Google's self-computed runs.
- Google reports 77.9% on DeepSWE v1.1 and calls it a new state of the art. Google computed that score itself, using a mini-swe agent harness.
- Google says Argon leads the Vals Index at 68.9%. The PDF says Vals AI supplied those results.
- Google says Argon ranks first on Zapier's AutomationBench at 51.3%. The PDF says that figure comes from Zapier's public leaderboard.
- Google reports 91.7% on LVBench, computed by Google without tools.
- Google says Argon ties for first on CWE-bench v1 at 68%. Its table gives GPT-6 Astra the same 68.0%.
The same table shows rivals ahead on several rows. GPT-6 Astra leads FrontierSWE v2 at 65.5% against Argon's 55.0%. Astra also leads Terminal-Bench Science 0.1 at 68.1% against 57.6%. On the OSWorld-2.0 offline subset, Astra scores 72.6% against Argon's 69.2%. Claude Opus 5.5 leads Terminal-bench 4.0 at 66.4% against 57.4%. Most rival figures come from other sources and setups, as the limits section explains.
Google also describes internal use. It says Argon agents freed over 300 TiB of memory across its data centers once rolled out. It says agents replaced 32K lines of SIMD code in a Rust port of libgav1. Google says the result runs 2.7x faster than that port. It says large rewrites still go through automated and manual auditing before production. These internal figures are Google's own and have no outside check.
On security, Google says Wiz used Argon through its Scan for Good program to find a critical flaw in healthcare software used by hospitals. Google says earlier frontier models had missed it. The post does not name the software.
No release date and uneven benchmark setups
Google gives no date for developer, enterprise or consumer access. The post gives no API model ID.
Google says it is "actively engaged in the U.S. government’s voluntary process for pre-release model access" while it widens access. The post does not name the agency or describe the testing.
The methodology PDF shows the comparisons do not share one setup:
- Google says rival scores mostly come from the providers' own reports or public leaderboards, at maximum reasoning settings where reported.
- Argon's OSWorld-2.0 score is the best of three runs.
- LVBench used 1 frame per second for Gemini, 800 frames for GPT-6 Astra, 300 for Fable 5.1 and 600 for Opus 5.5. Google cites API limits.
- Fable 5.1 has no Agent's Last Exam score, and Anthropic models have no OSWorld-2.0 offline score in the table.
Google lists four safeguard areas before broad release:
- It says Argon is designed to refuse harmful cyber and CBRN requests under its Frontier Safety Framework. Internal and external red teams tested those safeguards.
- It says Argon leads Gray Swan's Indirect Prompt Injection benchmark. The PDF says Gray Swan supplied that result.
- It says it is deploying monitors on Argon's chain of thought and actions that stop execution when necessary.
- It says it isolates and seals sandboxes before high-risk training or evaluations.
Apply to Fairwind or wait for the API
Security teams at governments, critical infrastructure operators and core technology platforms can apply through the Fairwind Program page. Plan for the MFA, team scoping and access tracking terms before you apply.
Everyone else has no access yet. When Argon reaches the API, budget at $2 and $10 per million tokens, and plan for $4 and $20 after the introductory period. The introductory rate matches the list price of GPT-6.1 Sol. Run both on your own harness before you switch.
Earlier gated cyber models from Google, OpenAI and Anthropic are compared in Cyber is now a product line: Daybreak, Fairwind, Mythos.
