Google has unveiled Gemini 4 Argon, its first model from the Gemini 4 generation and its most capable system to date. Announced on September 30 by Sundar Pichai, Google DeepMind and the Gemini developer team, Argon is initially rolling out only to a vetted group of cybersecurity defenders through a gated scheme called the Fairwind Program, with paid API access and Google AI Ultra subscriptions to follow at an unannounced date.
The staged launch makes Argon one of the clearest examples yet of a frontier lab putting cyber defence first when releasing a model with powerful coding and security capabilities.
What Google Announced
Argon is the first new flagship Gemini generation since Gemini 3 shipped last November. Google says it is designed for three main areas:
- Real-world software engineering, including long, multi-step coding tasks.
- Enterprise knowledge work such as legal and financial analysis.
- Cybersecurity defence, including finding and fixing vulnerabilities.
One headline specification is output length. Google says it is expanding Argon's output token limit to 1 million tokens, up from 64K on prior models, a change that could allow the model to generate entire codebases, long reports or large structured datasets in a single response.
Google AI leader Koray Kavukcuoglu wrote in the launch post that Argon is fundamentally changing how Google itself works and builds.
Benchmarks: Strong Claims, Not Yet Verified
According to Google's own figures, Argon sets new marks on several demanding evaluations:
- DeepSWE v1.1: 77.9%, ahead of Claude Opus 5.5 at 74.2% and GPT-6 Astra at 74.1%.
- LVBench (long video understanding): 91.7%.
- AutomationBench (end-to-end task automation): 51.3%, ranked first.
- CWE-bench v1 (vulnerability remediation): 68%, tied for first.
Google also says Argon beats GPT-6 Astra on 13 of 18 benchmarks it reported. However, every published score comes from Google or its pre-release cohort, and no independent lab has reproduced the results yet. Some coverage notes Argon trails Astra on terminal-driven coding tasks, and Bloomberg reported a less uniform internal picture among Google engineers about its real-world coding performance.
Pricing and Access
Google has published introductory API pricing well below many rival frontier models:
- $2 per million input tokens and $10 per million output tokens during the introductory period.
- Rising to $4 and $20 respectively afterwards.
- Cached input discounted by 95%.
General API model IDs and consumer rollout dates have not been published. Google says it is participating in the U.S. government's voluntary pre-release model access process, which allows federal evaluators to test frontier systems before wide release.
Why It Matters
Argon's launch strategy reflects a broader shift in how frontier labs handle models with significant offensive potential. By giving vetted defenders early access, Google aims to let security teams patch vulnerabilities before similar capabilities become widely available. Reports indicate defenders in the Fairwind Program receive the model without some of the standard cyber guardrails applied to consumer products, a trade-off that maximises defensive usefulness while restricting who can use it.
The approach also raises questions. Gated releases delay independent testing during the period when launch claims receive the most attention, making it harder for researchers and customers to validate benchmark results. With regulators, including the U.S. Federal Trade Commission, increasingly scrutinising AI safety practices, how labs balance capability, access and verification is becoming a competitive and policy issue in its own right.
For enterprises, the pricing is notable. If Argon's benchmark performance holds up under independent testing, its introductory cost structure could pressure rivals on price for coding and agentic workloads, particularly where long outputs and heavy caching are common.
The Competitive Landscape
Argon arrives in a crowded autumn for frontier models. Anthropic's Claude Opus 5.5 and OpenAI's GPT-6 family have both shipped in recent weeks, each emphasising agentic coding and computer use. Google's bet is that a combination of top benchmark results, a 1M-token output window and aggressive pricing will win developers once broader access opens.
For development teams planning ahead, the sensible move is to prepare evaluation suites now: a representative set of internal coding, document and security tasks that can be run against Argon the day API access opens, so that purchasing decisions rest on first-hand results rather than vendor leaderboards.
The decisive test will come when Argon reaches general developers and independent evaluators. Until then, Gemini 4 Argon is an impressive set of claims and a revealing case study in how frontier AI is now released: carefully, in stages, and with cybersecurity at the front of the queue.
