Google announced Gemini 4 Argon on September 30, 2026. Google pitches it at long-horizon reasoning, software engineering, enterprise knowledge work in law and finance, cyber defence and multimodal understanding. The most unusual spec is output length: Google says it is "significantly expanding the model's output token limit to an industry-leading 1M tokens, up from the previous 64K tokens."
Google's headline numbers:
- DeepSWE v1.1: 77.9% on long software-engineering tasks.
- AutomationBench: first place at 51.3%. Vals Index: first place, which Implicator puts at 68.9%.
- CWE-bench v1: 68%, tied for first on fixing vulnerabilities. LVBench: 91.7% on long-video understanding.
- Introductory pricing: $2 per million input tokens and $10 per million output tokens, with a 95% discount on cached input. Standard pricing afterwards is $4 and $20.
The rollout comes before general availability. Argon is going first "to a set of trusted cyber defenders through our Fairwind Program." The Next Web reports that those partners get the model without its cyber guardrails, and that Wiz used it to find a critical vulnerability in healthcare software that earlier frontier models had missed. Paid API customers and Google AI Ultra subscribers come next, with no date given.
The launch also brought a credibility problem. Bloomberg reported, as summarised by Implicator and The Next Web, that some Google employees who have used Argon say it does worse on real coding work, especially front-end design, than its scores suggest. Two insiders reportedly blamed "benchmaxxing." Google disputed the characterisation. Gemini product head Tulsee Doshi said "many" Googlers rely on it "for their hardest coding and research problems." Implicator notes that the model ties GPT-6 Astra and Claude Fable 5.1 at 53 on the Artificial Analysis Intelligence Index, while trailing competitors at 57% on Terminal Bench 4.
Why it matters: Google is releasing its flagship to vetted defenders, with cyber guardrails removed for them, before releasing it to the public. It is a security-first rollout ahead of the public. For buyers, the disagreement between benchmark scores and internal feedback is a reminder to test on your own repositories before switching. Watch for a general-availability date, independent agentic-coding results, and whether a 1M-token output limit holds up in practice or mainly serves long code and document generation.
Gemini 4 Argon leads several Google-reported benchmarks at $2/$10 introductory pricing with a 1M-token output limit, but it is reaching cyber defenders before the public, and reported staff doubts about its real coding performance make in-house evaluation essential before switching.