Science & Tech · United States of America
Google Launches Gemini 4 Argon, Restricts Early Access to Cybersecurity Partners
Google's new flagship AI model claims leading benchmark scores in coding, cybersecurity and professional tasks, but the company says it will first go only to trusted cyber-defense partners while safety reviews continue.
Google announced Gemini 4 Argon, its most advanced AI model, restricting initial access to cybersecurity defense partners while safety checks continue.
- Argon is Google's first flagship model since Gemini 3, released Sept 30
- Initial access limited to cybersecurity defenders via the Fairwind program
- Output limit raised to 1 million tokens; pricing set at $2/$10 per million tokens
- Google says Argon leads or ties rivals on many benchmarks
- Google is in the U.S. government's voluntary pre-release safety review
What's new
- Google unveiled Gemini 4 Argon, its first flagship model since Gemini 3 nearly a year ago
- Initial access limited to cybersecurity defense organizations through Google's Fairwind program
- Output token limit raised from 64,000 to 1 million tokens
- Google reports Argon leads or ties rivals from OpenAI and Anthropic on several benchmarks
Google said on Wednesday that it had begun rolling out Gemini 4 Argon, which it calls its most advanced artificial intelligence model, giving initial access only to organizations focused on cybersecurity defense while additional safety evaluations continue.157
A limited rollout
Google says Fairwind will give trusted cybersecurity defenders the earliest access to Argon, with the model becoming available afterward to subscribers of Google AI Ultra and to customers paying for API access. The company did not provide a timeline for public availability.59108
For its introductory pricing, Google is charging $2 for every million input tokens processed by Argon and $10 for every million tokens it outputs.934
Benchmark claims
Google claims Argon outpaces rival systems built by OpenAI and Anthropic on numerous benchmark tests, citing a DeepSWE v1.1 coding score of 77.9 percent versus 74.2 percent for Anthropic's Claude Opus 5.5 and 74.1 percent for OpenAI's GPT-6 Astra. Google also reports Argon scoring 68.9 percent on the Vals Index, a benchmark covering finance, legal and professional tasks, compared with 67.0 percent for Opus 5.5.46
On Zapier's AutomationBench, Google reports Argon at 51.3 percent against 42.5 percent for Opus 5.5. Anthropic's model still leads on Terminal-Bench 4.0, according to Google's figures, at 66.4 percent versus Argon's 57.4 percent, though Google noted that Artificial Analysis separately measured Opus 5.5 at only 59.6 percent on that test. On the CWE-bench v1 cybersecurity test, Argon matched Astra's score of 68 percent, edging out Opus 5.5's 67 percent.4
Internal use and doubts
Google says thousands of its employees already use Argon internally for coding, research and writing. The company reports that Argon-driven memory optimizations have freed more than 300 tebibytes of memory across its data centers without additional hardware purchases, and that Argon agents migrated more than 800,000 lines of code in the Fuchsia OS Zircon kernel. A video decoder optimized by Argon agents is reported to run 2.7 times faster than the previous version written in Rust.46
Axios reported that some Google staff found Argon's performance lacking, a claim Google disputed. Bloomberg separately reported internal doubts about the model's benchmark figures, according to Trending Topics.104
Safety review
Google is participating in the United States government's voluntary process for pre-release access to new AI models, according to multiple outlets. A day before the Argon announcement, Google chief executive Sundar Pichai co-signed a voluntary safety accord with President Trump alongside other major technology companies.7823
Google said it plans to strengthen critical safeguards before a broader public rollout of Argon.75
Why it matters
Argon intensifies competition among the leading AI developers at a moment when European businesses are weighing which models to adopt for coding, finance and legal work.27
Related stories
Version history
- version 1 ·
