Gemini 4 Argon is Google’s new frontier model, and on paper it is the best coding model in the world right now. Google says it scores 77.9% on DeepSWE v1.1, ahead of Claude Opus 5.5 and GPT-6 Astra, at an introductory price of $2 per million input tokens.

There is a catch the launch headlines bury. Unless you work for a vetted cybersecurity organisation, you cannot use it yet.

I spent the last two days reading Google’s launch post, the benchmark tables, and the early coverage. This is what is real, what is marketing, and what it changes for developers who are choosing a model this month.

What Is Gemini 4 Argon?

Gemini 4 Argon is Google DeepMind’s frontier model, announced on September 30, 2026, and positioned as the strongest model available for agentic coding and cybersecurity. Google describes it as the start of a “next era of frontier intelligence”, which is the standard launch-post register, so I am going to ignore the adjectives and look at the numbers.

The headline specs from Google’s announcement:

  • Output limit: 1 million tokens, up from 64K on previous Gemini models
  • Context window: 2 million tokens, according to Yahoo Finance’s launch coverage (Google’s own post does not state it)
  • Introductory pricing: $2 input / $10 output per million tokens
  • Standard pricing: $4 input / $20 output after the introductory period
  • Cached input: 95% off the input price

The output limit is the spec I find most interesting. A 1 million token output means one response can contain a rewritten module, test suite and migration notes, without the agent having to chunk the work into dozens of calls. Google’s own example shows why it matters, which I cover below.

The short version: Argon is a coding and security model first, and Google priced it to compete directly with Anthropic.

Gemini 4 Argon Benchmarks: The Coding Numbers

Here are the scores Google published, alongside the comparison figures reported in Yahoo Finance’s coverage of the launch:

BenchmarkGemini 4 ArgonClaude Opus 5.5GPT-6 Astra
DeepSWE v1.1 (agentic coding)77.9%74.2%74.1%
CWE-bench v1 (vulnerability fixing)68% (tied 1st)——
AutomationBench51.3% (1st)——
LVBench (video understanding)91.7%——

A 3.7-point lead on DeepSWE is real but not dramatic. That is roughly the gap between Opus 5 and Opus 5.5, a generation apart within one vendor.

The less flattering number did not make Google’s headline. Several third-party launch breakdowns report Argon at 55.0% on FrontierSWE v2, a harder benchmark where the lead is much less clear. Anthropic’s own table puts Claude Opus 5.5 at 54.4% on FrontierCode v1.1, a different benchmark, so I would not read too much into the comparison. It does show that “state of the art” depends on which test you pick.

That leads to the caveat that applies to this entire section. Every Gemini 4 Argon benchmark is vendor-reported, and no independent lab has verified any of them yet. Yahoo Finance’s analysis says so directly. Anthropic’s and OpenAI’s numbers have the same problem. Each lab runs the others’ models through its own harness, at effort settings it chooses.

My rule for frontier launches: treat any lead under 5 points as a tie until Artificial Analysis, METR or an equivalent independent group re-runs it. By that rule, Argon is tied with Opus 5.5 and Astra on agentic coding. It is not yet a clear winner.

The Rust Migration Claim Is the Real Story

The benchmark table is the least interesting part of the launch. The part that made me stop scrolling is how Google says it has been using Argon internally.

According to the launch post, Argon has been migrating C and C++ codebases to Rust across Google. These range from tens of thousands of lines in core libraries like re2 and libgav1 up to more than 800,000 lines for the Fuchsia Zircon kernel. On libgav1, Google says the Argon-produced code ran 2.7x faster than the existing Rust port.

This is a different kind of claim from a benchmark score. A kernel migration of that size has to compile, pass existing tests, and survive review from engineers who know the original code very well. If the claim holds up, it is the strongest public evidence so far that frontier models can handle large, real legacy migrations, not just isolated GitHub issues.

It also explains the 1 million token output limit. You cannot rewrite a kernel subsystem 64K tokens at a time without the agent losing track of what it already changed.

Anthropic is making a similar argument. Its Opus 5.5 launch cites a 680,000-line migration completed in under a day, and a HAProxy C-to-Rust translation. Large-scale migration is now the benchmark the labs actually compete on, even if no leaderboard tracks it yet.

For teams sitting on old C/C++ or Java code, this capability matters more than any leaderboard number.

Who Can Use Gemini 4 Argon Today?

As of October 2, 2026, Gemini 4 Argon is available only to trusted cyber defenders through Google’s Fairwind Program. Yahoo Finance reports more than 650 organisations in that group, including government agencies and security vendors such as CrowdStrike and Palo Alto Networks.

Google says the next wave will be paid Gemini API customers, followed by Google AI Ultra subscribers. Its only timeline is “as soon as possible”. There is no date for AI Studio, Vertex AI, or the consumer Gemini app.

The staged rollout has a clear reason. Google says Argon can autonomously find, validate and patch critical software vulnerabilities, and the Fairwind version runs without cyber guardrails for vetted defenders. In a partnership with Wiz, Argon found a critical vulnerability that exposed sensitive personal data across healthcare software used by hospitals worldwide.

A model that finds a vulnerability like that for a defender can find it just as easily for an attacker. Letting defenders patch before anyone else gets access is the responsible order. It is also very frustrating if you just want to point the model at your monorepo.

Anthropic took a similar route with Opus 5.5. It runs a three-tier Cyber Verification Program and keeps general-access safeguards in place. Frontier models are now strong enough at security work that every lab is gating that capability in some way.

The practical conclusion: do not plan a Q4 project around Gemini 4 Argon API access. If it arrives, treat it as a bonus.

Gemini 4 Argon Pricing vs Claude Opus 5.5 and GPT-6

Pricing is where Argon could change developer behaviour, once it is generally available.

ModelInput ($/M)Output ($/M)Cache discountAvailable now?
Gemini 4 Argon (intro)$2$1095% offNo, Fairwind only
Gemini 4 Argon (standard)$4$2095% offNo
Claude Opus 5.5$4$20$0.20 cache readsYes
Claude Sonnet 5.5$2$10$0.20 cache readsYes
GPT-6 Astra$10$50$1.00 cached inputYes
GPT-6 Sol$2$10up to 90% offYes

(Pricing from Google, Anthropic and OpenAI launch materials and MarkTechPost’s coverage, September–October 2026.)

Three things stand out.

First, Argon’s standard price is exactly the same as Claude Opus 5.5, at $4 in and $20 out. That is very unlikely to be a coincidence. Google is pricing for the same buyer Anthropic is.

Second, the introductory price matches Claude Sonnet 5.5 and GPT-6 Sol at $2 / $10. For the intro window, Google is offering its frontier model at the mid-tier price of its competitors. If you get access, that is a very good deal.

Third, GPT-6 Astra at $10 / $50 now looks expensive. OpenAI’s flagship costs 2.5x Argon’s standard input price and 5x its intro price. Astra has to win on quality by a wide margin to justify that, and the vendor-reported numbers say it does not.

The 95% cache discount matters more than the headline rates for agentic coding. In a typical coding agent session, most input tokens are the same files and instructions re-sent every turn. At 95% off, cached reads on standard pricing cost $0.20 per million, the same as Claude Opus 5.5’s cache read price. In cost per agentic task, the two will probably be very close.

Bottom line on pricing: Argon at standard rates costs the same as Opus 5.5, and the intro rate undercuts it by half.

What Gemini 4 Argon Changes for Developers Right Now

Here is my honest read on what to do with this launch this month.

If you use Claude Code or Codex: nothing changes yet. Claude Opus 5.5 and GPT-6 Sol are available today, they are within a few points of Argon on Google’s own numbers, and they work in tools you already use. My comparison of Claude Code, Cursor and Codex still applies. For the model choice itself, see my Claude Opus 5.5 vs GPT-6 coding comparison.

If you are on Google’s stack: watch for the paid API rollout and test Argon on your hardest migration task as soon as you can. The Rust migration claim is the one to verify with your own code. Google’s own agent tooling, which I looked at in the Antigravity deep dive, is where Argon will probably appear first for developers.

If you work in security: apply to Fairwind if your organisation qualifies. A model that can find and patch vulnerabilities without guardrails is the most useful version of this release. The gap between AI-written code and AI-audited code is a theme I keep returning to, and the AI-generated code security numbers are why. Argon could finally close part of that gap.

If you are choosing a model on benchmarks: wait for independent results. A 3.7-point vendor-reported lead tells you that the three labs are close. It does not tell you which model will work best on your codebase.

The Verdict on Gemini 4 Argon

Gemini 4 Argon is a serious frontier model with an unusual launch. The strongest evidence in Google’s announcement is the 800K-line kernel migration, not the benchmark table. The most important number for developers is the $4 / $20 standard price, which puts Google in direct competition with Anthropic.

The gated rollout is the right call for a model with these security capabilities. It also means that, in October 2026, Gemini 4 Argon is mostly something to read about, not something to build on.

I will update this post when independent benchmarks land and when the paid API opens up.