← All posts / Models

One Million Tokens of Thinking: Google Announces Gemini 4 Argon, Its New Frontier Model — but Almost No One Can Use It Yet

Google DeepMind unveils Gemini 4 Argon: state-of-the-art on DeepSWE and the Vals Index, a 1M-token output ceiling, $2/$10 pricing — yet initially restricted to trusted cyber defenders via the Fairwind Program.

One Million Tokens of Thinking: Google Announces Gemini 4 Argon, Its New Frontier Model — but Almost No One Can Use It Yet

At 8 PM UTC on September 30, 2026, Google DeepMind pulled back the curtain on Gemini 4 Argon — its next frontier model, and the most consequential AI release of the fall season. The twist: almost nobody outside a hand-picked cohort of cybersecurity defenders and Google’s own engineers can touch it yet. In an era when frontier launches usually mean immediate API access and a livestreamed demo marathon, Google has chosen the opposite path — a deliberately staged rollout that says as much about the politics of AI in 2026 as it does about the model itself.

What Argon Is

Google describes Gemini 4 Argon as a frontier model “built to sustain deep reasoning across complex, long-horizon workflows.” The positioning around three pillars — real-world software engineering, enterprise knowledge work like legal and finance, and cybersecurity defense — marks a clear strategic bet: Argon is aimed not at chatbot trivia but at sustained professional work, the kind that takes hours or days for a human team and produces measurable economic value.

The headline technical spec is the output token limit. Argon expands the ceiling from the previous 64K tokens to an industry-leading 1 million tokens per response. That number matters more than it might first appear. When a model can generate hundreds of thousands of tokens in a single reasoning trajectory, it can attempt genuinely long-horizon tasks — full-scale codebase migrations, end-to-end security audits, deep multi-document legal analysis — without the context-switching and state-loss that break up agentic workflows today. It converts the “reasoning budget” from a snack into a buffet.

Pricing, when Argon does reach the API, undercuts rivals aggressively: an introductory $2 per million input tokens and $10 per million output tokens, with cached input at 95% off. That underprices OpenAI’s GPT-5.5 ($5/$30) by a wide margin and signals that Google intends to compete on cost-per-task as much as on benchmark scores.

The Numbers

Across the 18 benchmarks Google disclosed, Argon leads outright on 12 and ties for first on one, according to VentureBeat’s analysis — with GPT-6 Astra leading on three and tying on one. Highlights:

  • DeepSWE v1.1: 77.9% — a new state of the art on the benchmark for real-world, long-horizon software engineering tasks, the discipline where AI is expected to deliver its most direct economic payoff.
  • Vals Index: leading — the index that measures economic impact across finance, coding, legal, and tax work, with each sector weighted by its contribution to U.S. GDP. Argon is the top model, with similarly leading scores on Vals Finance Agent v2 (multi-step financial research) and Harvey’s Legal Agent Benchmark (legal research and drafting).
  • AutomationBench: 51.3%, ranked #1 — Zapier’s benchmark for end-to-end execution across core business functions.
  • LVBench: 91.7% — state of the art on long-video understanding, reflecting Argon’s strength when knowledge work demands visual comprehension.
  • CWE-bench v1: 68%, tied for first — on remediating security vulnerabilities, building on 3.8 Flash Cyber’s earlier frontier results.
  • Gray Swan IPI: 0.7% attack success rate — leading the pack on robustness against indirect prompt injection, a critical defense for agentic deployments (compared with 1.0% for the nearest competitor).

The New Stack’s early assessment adds a dose of nuance: Argon’s knowledge-work wins are decisive, but its coding scores are “mixed” relative to OpenAI and Anthropic on some measures. DeepSWE’s 77.9% is state of the art, but the leaked “Gemini 4 Pro” checkpoint scores circulating days earlier (88.7% DeepSWE) suggest Google may be holding back the full-capability variant — or that those numbers were never real.

Argon at Work Inside Google

The most striking section of the announcement is what Argon is already doing internally:

  • Quantum computing: Argon is helping researchers optimize the spacetime resources (qubits × gates) of subroutine bottlenecks. In one example it beat the published baseline by 40% in minutes.
  • Data-center efficiency: A team of Argon agents analyzed fleet-wide profiling telemetry and autonomously applied memory optimizations across Google’s data centers, freeing over 300 TiB of memory, with an estimated 500 TiB to 1 PiB in total savings.
  • Rust migrations: Argon agents are porting C/C++ codebases to Rust — from tens of thousands of lines in core libraries like re2 and libgav1 up to 800K+ lines for the Fuchsia Zircon kernel. For libgav1, agents replaced 32K lines of SIMD code through profile-guided experiments, producing a memory-safe video decoder that runs 2.7× faster than the previous Rust port with identical output.

These are not marketing demos; they are infrastructure-scale deployments with auditable results. The 1M-token output limit is precisely what makes such single-trajectory rewrites feasible.

Why Cyber Defenders Get It First

Argon’s first external users are “trusted cyber defenders” enrolled in Google’s Fairwind Program — the initiative announced in early September to give defenders access to advanced Gemini models for autonomously finding and fixing vulnerabilities. For this cohort, and for Google’s internal teams, Argon ships without cyber guardrails, unlocking its full frontier-level defensive capabilities.

The early proof point is tangible: Wiz is already using Argon in its Scan for Good initiative, which protects critical public infrastructure for free. In an early deployment, the model uncovered a critical vulnerability exposing sensitive personal information across healthcare software used by hospitals worldwide — a severe risk that previous frontier models had missed. On Wiz’s internal black-box penetration-testing benchmark, Argon outperformed 3.8 Flash Cyber in discovering attack surfaces, identifying vulnerabilities, and producing proof-of-concept evidence.

The defensive-first rollout inverts the usual offense-heavy narrative of AI security. Instead of asking “what can this model break,” Google is demonstrating what it can protect — and building a public trust case before general release.

Safeguards and the Staged Rollout

Google is explicit that “safely releasing frontier capabilities at this level requires a phased approach,” noting it is “actively engaged in the U.S. government’s voluntary process for pre-release model access.” That phrase ties Argon to the governance framework the White House built this year — a voluntary pre-release review channel that frontier labs have adopted in the absence of binding regulation.

Four safeguard pillars accompany the launch: refusing harmful cyber and CBRN requests while preserving legitimate dual-use research (per the Frontier Safety Framework), with internal-activation monitoring to spot misuse; leading prompt-injection robustness via automated red teaming and adversarial training; chain-of-thought and action monitoring that halts execution when the model steps beyond user intent; and hardened, sealed sandbox environments for high-risk evaluations, in line with Google’s agent control roadmap. Notably, Google used a similar monitoring system during training runs and is publicly urging the industry to preserve reasoning transparency “in these pivotal moments of increased capabilities.”

What It Means

Three takeaways stand out. First, the benchmark lead has genuinely rotated back to Google — at least until OpenAI or Anthropic answers — and the margin on economically-weighted evaluations like the Vals Index makes the claim concrete rather than rhetorical. Second, the 1M-token output window changes the anatomy of agentic work: tasks that previously required orchestration frameworks to stitch together can now run as single trajectories. Third, the distribution strategy is the story inside the story. By gating Argon behind Fairwind, government pre-release engagement, and a paid-API-first expansion, Google is effectively treating frontier capability as a controlled substance — normalizing the idea that the most powerful models launch quietly, to vetted users, under supervision.

For developers, enterprises, and consumers, Argon arrives “soon,” starting with paid API customers and Google AI Ultra subscribers. Whenever that happens, the $2/$10 price point will apply immediate pressure on every frontier pricing sheet in the industry. Until then, the rest of us watch from outside the glass — while a model that can think for a million tokens goes to work on hospital software, quantum subroutines, and Google’s own kernel.