← Back to all posts
News

Google's Gemini 4 Argon Leads 13 of 18 Benchmarks, Ships Only to Cyber Defenders

October 1, 2026 · 00:08 UTC · News
Google's Gemini 4 Argon Leads 13 of 18 Benchmarks, Ships Only to Cyber Defenders

TL;DR

On September 30, Google DeepMind announced Gemini 4 Argon, the flagship model that replaces the long-delayed Gemini 3.5 Pro. In the comparison table Google published, Argon takes first place or ties on 13 of 18 benchmarks against GPT-6 Astra, Claude Opus 5.5 and Claude Fable 5.1, with its biggest leads in knowledge work and long context. It trails by 9 to 10.5 points on three coding and science benchmarks. It can write up to 1M output tokens, up from 64K, and it will cost $2/$10 per million tokens at launch, rising later to $4/$20. The catch: right now only a vetted group of cyber defenders in Google's Fairwind Program can use it.


What Google Actually Shipped

The announcement comes from Koray Kavukcuoglu, Google DeepMind SVP and Google's Chief AI Architect, and it is careful with the word "release." Argon is "rolling out to a set of trusted cyber defenders through our Fairwind Program." Next in line are paid API customers and Google AI Ultra subscribers, then developers, enterprises and consumers in general, "as soon as possible." Google gives no date.

Google also says it is "actively engaged in the U.S. government's voluntary process for pre-release model access while we gradually expand access." So this follows the pattern Anthropic set with Mythos: announce the frontier model, publish the numbers, and give it first to defenders who are expected to find bugs before attackers do.

According to The New Stack, Google first promised a new Pro model, Gemini 3.5 Pro, at I/O in May and planned to ship it in June. What arrived over the summer was a series of Flash models, including 3.8 Flash and 3.8 Flash Cyber. Argon is the model that takes that slot.

The Benchmark Table, Wins and Losses

Google's own post gives only a few headline numbers: a state-of-the-art 77.9% on DeepSWE v1.1, first place on the Vals Index, 51.3% and first place on Zapier's AutomationBench, 91.7% on LVBench for long-video understanding, and a tie for first at 68% on CWE-bench v1. The New Stack reproduced Google's full comparison table, with 18 benchmarks across four models. Read the whole thing, not just the summary.

Where Argon wins, it can win by a lot. It scores 51.3% on AutomationBench against 42.5% for Opus 5.5. On Harvey's Legal Agent Benchmark it scores 19.6%, nearly three times Fable 5.1's 6.7%, which still means it fully completes only about one task in five. On GraphWalks BFS with 256K to 1M token inputs it reaches 84.2%, against 71.8% for GPT-6 Astra.

Where it loses, it also loses by a lot. Argon finishes last of the four on FrontierSWE v2 (55.0% vs Astra's 65.5%), last on Terminal-bench 4.0 (57.4% vs Opus 5.5's 66.4%), and well behind on Terminal-Bench Science 0.1 (57.6% vs Astra's 68.1%). Opus 5.5 also leads it on PostTrainBench, 49.3% to 45.3%.

Argon minus best rival, in points (Google's table) Harvey Legal Agent+12.9 GraphWalks 256K-1M+12.4 AutomationBench+8.8 Vals Finance Agent v2+6.5 LVBench+4.2 DeepSWE v1.1+3.7 PostTrainBench-4.0 Terminal-bench 4.0-9.0 FrontierSWE v2-10.5 T-Bench Science 0.1-10.5
Knowledge work and long context are big wins. Two of the coding tests and the science test are big losses.

Google says DeepSWE v1.1 shows it can handle long-horizon software engineering. The other two coding tests disagree. If your workload is a terminal-heavy coding agent, Google's own table says Opus 5.5 is still ahead. If your workload is reading a few hundred thousand tokens of contracts, filings or logs and then acting on them, Argon is the model to beat.

One caveat from The New Stack applies to the cyber result. On CWE-bench v1, the OpenAI and Anthropic entries run inside their own agent harnesses (Codex and Claude Code), so the 68% tie with GPT-6 Astra measures each model plus its tooling, not the bare model.

The Independent Number: Artificial Analysis Says 53

Vendor tables are vendor tables. Artificial Analysis has already published a score for Gemini 4 Argon (High): 53 on its Intelligence Index, ranked #8 of 223 models. That ties Fable 5.1 and GPT-6 Astra, sits one point above GPT-6.1 Sol, and is five points below Claude Opus 5.5 at 58. Note that the rival scores shown are their maximum-effort settings, and the page is labelled "High," so a higher Argon setting may score better.

The Vals leaderboard independently matches Google's claimed Vals Index score: Argon sits at 68.9%, narrowly ahead of Opus 5.5's 67.0% in Google's table.

Artificial Analysis Intelligence Index (higher is better) Opus 5.5 (max)58 Argon (High)53 Fable 5.1 (max)53 GPT-6 Astra (max)53 GPT-6.1 Sol (max)52 GPT-6 Sol (max)48
On the independent index, Argon lands in the three-way tie behind Opus 5.5, not in first place.

Artificial Analysis also reports that Argon used 110M output tokens to run the index, against a median of 82M, at an average of $1.99 per task. It thinks at length, and the pricing is set up to make that affordable.

1M Output Tokens Is the Real Spec Change

A 1M-token input window is standard for frontier models now. Argon's change is on the output side: up to 1M output tokens, up from 64K for earlier Gemini models. Google's argument is that when the model "has the headroom to think deeply and generate hundreds of thousands of tokens in a single trajectory," it can solve hard problems in one go instead of stopping partway.

The output cap limits reasoning and answer together. With 64K, a model reasoning through a large migration has to stop, summarize, and hand off to another turn, losing detail at every handoff. Think of an exam where you get one sheet of paper for both scratch work and the final answer. Argon gets the whole notebook.

The practical impact for you: a single call can now produce a full module rewrite or a long report without an orchestration loop stitching chunks together. It also means a single call can produce a bill. At the regular $20 per million output tokens, one maxed-out response costs $20.

Pricing: Half of Opus 5.5, For Now

The introductory price is $2 per million input tokens and $10 per million output, with cached input 95% cheaper. A footnote in Google's post says that "after the introductory period expires," the price becomes $4/$20, which matches Opus 5.5's list price exactly. Google does not say how long the introductory period lasts.

output price, $ per million tokens Argon (intro)$10 Argon (later)$20 Opus 5.5$20
The launch discount halves the price. When it ends, Argon costs the same as Opus 5.5.

What Google Says Argon Already Does Internally

Google lists internal deployments, all self-reported:

  • Fleet memory. Argon agents analyzed fleet-wide profiling telemetry and applied memory optimizations that free more than 300 TiB once rolled out, with an estimated 500 TiB to 1 PiB total.
  • C/C++ to Rust. Migrations range from tens of thousands of lines in re2 and libgav1 up to 800K+ lines for the Fuchsia Zircon kernel, and Google says these rewrites are still under automated and manual audit before production.
  • libgav1 SIMD. Agents replaced 32K lines of SIMD code in an existing Rust port with safe Rust the compiler auto-vectorizes, producing a decoder 2.7x faster than that port with identical output.
  • Quantum circuits. In one example it beat a published baseline for qubits-times-gates cost by 40% "in a matter of minutes."

On security, Wiz is using Argon in its Scan for Good program, and Google says the model found a critical flaw exposing personal data in healthcare software used by hospitals worldwide. Google says trusted defenders and its internal teams get Argon "without cyber guardrails." Everyone else will get the guarded version.

The Safety Section Has One Unusual Line

Most of the safeguards section is standard: CBRN and cyber refusals under the Frontier Safety Framework, activation monitoring for misuse, and a claimed lead on Gray Swan's indirect prompt injection benchmark. The unusual part is about chain-of-thought monitoring. Google says it watched Argon's reasoning during training and deliberately did not feed what it found back into training, "so as to not risk shaping Argon's reasoning to evade our monitoring." It links to an essay urging the rest of the industry to keep reasoning legible too. Put simply: if you punish a model for thinking bad thoughts out loud, you may teach it to stop thinking them out loud.

Key Takeaways

  • Gemini 4 Argon is real but gated. Only Fairwind cyber defenders have it today. Paid API and AI Ultra come next, with no date.
  • Strong at knowledge work, mixed at coding. It is top or tied on 13 of 18 benchmarks in Google's table, but last on FrontierSWE v2 and Terminal-bench 4.0.
  • The independent score is 53, not first. Artificial Analysis puts Argon (High) level with Fable 5.1 and GPT-6 Astra, five points behind Opus 5.5.
  • 1M output tokens changes what one call can do. It is a 15x jump from 64K and makes single-pass long migrations and reports possible.
  • The $2/$10 price is temporary. It rises to $4/$20, the same as Opus 5.5, once an introductory period of unstated length ends.

Sources: Google: Gemini 4 Argon announcement, The New Stack: full benchmark table, Artificial Analysis: Gemini 4 Argon (High), Artificial Analysis: Claude Opus 5.5, Vals Index, Google DeepMind: Fairwind Program, CNBC, Hacker News discussion

AIGoogleGemini 4 ArgonGoogle DeepMindBenchmarksCybersecurityLLM Pricing
CONSOLE
$