Google Unveils Gemini 4 Argon, Claims Benchmark Lead but Limits Access to Cyber Defenders
Google's new top-tier Gemini 4 model claims leads over OpenAI's GPT-6 Astra and Anthropic's Opus on most disclosed benchmarks, but it is initially available only to vetted defenders through the Fairwind Program.
Google on September 30 announced Gemini 4 Argon, the first model in its Gemini 4 generation, describing it as a frontier model for real-world software engineering, enterprise knowledge work and cyber defense. The announcement came in a blog post from Koray Kavukcuoglu, Google's chief AI architect. The model is not yet broadly available.
Access is restricted at launch
Argon is initially being offered to "trusted cyber defenders" through Google's Fairwind Program, the early-access initiative the company launched in early September alongside Gemini 3.8 Flash Cyber. Google says it is also taking part in the U.S. government's voluntary pre-release model access process, and that broader availability will follow, starting with paid API customers and Google AI Ultra subscribers. It gave no date. According to TechCrunch, Google staff are already using the model internally for debugging and codebase migration.
Google framed the staged rollout around safeguards in four areas: misuse defense, prompt-injection robustness, misalignment monitoring and hardened sandboxed environments.
What Google claims
Google says Argon scores significantly higher than OpenAI's GPT-6 Astra and Anthropic's Fable and Opus models across a range of benchmarks. Figures published by Google include 77.9% on DeepSWE v1.1, 51.3% on AutomationBench, 91.7% on LVBench for long-video understanding, and a tie for first at 68% on CWE-bench v1, a vulnerability-remediation test. It also cites leadership on the Vals Index, which covers finance, coding, legal and tax work. The model supports up to 1 million output tokens, up from 64,000 on previous Gemini models.
VentureBeat's tally of Google's disclosures found Argon leading or tied on 13 of 18 benchmarks, with examples including 77.9% on DeepSWE v1.1 versus 74.2% for Claude Opus 5.5, and 51.3% on AutomationBench versus 42.5%. GPT-6 Astra still leads on some tests, notably FrontierSWE v2 and Terminal-Bench Science. These are vendor-reported results; independent testing is limited while access is restricted.
Pricing
Google lists introductory API pricing of $2 per million input tokens and $10 per million output tokens, moving to $4 and $20 afterward, with a 95% discount on cached input tokens. VentureBeat notes that undercuts GPT-6 Astra's $10 and $50 and matches Claude Opus 5.5 once the introductory period ends. Because most developers cannot yet call the model, whether the benchmark gains and pricing hold up on production workloads remains untested.
Why it matters
Argon follows a pattern set by OpenAI's Astra and Anthropic's Mythos models: the most capable cyber-relevant systems reach vetted defenders first, with general release gated behind safeguards and government review. For Gemini users, the practical change is still ahead; the headline is that Google is again competing for the top of the leaderboard.
Sources
AI-assisted reporting, overseen by the AgentsAI team. Spotted an error? Let us know.
Related agents
More ai news
EU Set to Propose Barring Under-15s From AI Chatbots and Social Media in 'Kids Act'
The European Commission is preparing to unveil an EU Kids Act that would bar unsupervised access to AI chatbots, social media, video platforms and online games for under-15s, with tiered rules and mandatory age verification for 13-14 year-olds.
Amodei's 'Pace the Frontier' Plan Draws Same-Day Backing From OpenAI, DeepMind and xAI
Anthropic CEO Dario Amodei published an essay arguing frontier AI labs should deliberately slow capability gains, and within hours Sam Altman, Demis Hassabis and Elon Musk publicly endorsed the idea, with Microsoft's Satya Nadella following a day later.
Positron Raises $875M to Build an HBM-Free AI Inference Chip
Chip startup Positron closed an $875 million Series C at a $5 billion post-money valuation to fund its Asimov inference accelerator, which pairs its compute architecture with up to 2,304GB of commodity LPDDR5X memory instead of scarce high-bandwidth memory.
DeepSeek Releases V4.1 Flash, Cuts API Prices and Sets End Date for V4 Pro
DeepSeek officially released V4.1 Flash, a cheaper and faster multimodal successor to V4 Pro with a 1-million-token context window, and said it will reroute all V4 Pro API traffic to the new model from September 14.