
What Is Gemini 4 Argon? Google's New Model Beats Claude Fable 5.1 on 15 of 17 Comparisons and Outscores GPT-6 Astra
Introduction
On September 30, 2026, Google announced Gemini 4 Argon, the first model in its Gemini 4 generation. CEO Sundar Pichai posted on X that there was "lots of discussion out there about our next model," so he wanted to give an early look, and shared a table comparing Argon with Claude Fable 5.1, GPT-6 Astra, and Claude Opus 5.5.
Key takeaways
- In Google's official table, Argon beats Claude Fable 5.1 on 15 of the 17 comparisons where both have scores. Fable 5.1 wins only FrontierSWE v2 and Terminal-bench 4.0
- Of the table's 19 rows, Argon leads outright on 13. GPT-6 Astra leads on three, Claude Opus 5.5 on two, and Argon and Astra tie on one
- Introductory API pricing is one-fifth of Fable 5.1 and Astra, but for now only cyber defenders and some trusted testers have access
A leak claiming that "Gemini 4 Pro beats Astra and Fable" circulated before the launch. Using the official table, this article shows where Argon pulls ahead of Fable 5.1 and Astra, where it trails, and how the leak compared.
What is Gemini 4 Argon?
Gemini 4 Argon is Google's new frontier model, built to sustain long, complex workflows in a single run. In the official blog post, Google DeepMind's Koray Kavukcuoglu focuses on three areas: real-world software engineering, enterprise knowledge work such as legal and finance, and cybersecurity defense.
The key points:
- Output limit: Raised from 64,000 to 1 million tokens per response, so long reasoning chains and large code migrations can run without handing control back to a human
- Pricing: $2 per million input tokens and $10 per million output tokens during an introductory period, with cached input 95% off. After that, $4 and $20
- Availability: First to cyber defenders in Google's Fairwind Program, while Google takes part in the US government's voluntary pre-release access process. Paid API customers and Google AI Ultra subscribers come next
Thousands of Googlers already use Argon internally. Google says Argon agents replaced 32,000 lines of SIMD code in the libgav1 video decoder with safe Rust, making it 2.7 times faster than the existing Rust port. Argon agents also found memory optimizations expected to free more than 300 TiB across Google's data centers.
This is Google's first new flagship in a while. After Gemini 3 in November 2025, Google cancelled Gemini 3.5 Pro and kept iterating on its lighter Flash line.
Where does it beat Claude Fable 5.1?
Of the 17 comparisons in the official table where both Argon and Fable 5.1 have scores, Argon wins 15. The widest gaps are in business automation, science, long context, and chart understanding.
| Area | Benchmark | Gemini 4 Argon | Claude Fable 5.1 | Gap |
|---|---|---|---|---|
| Knowledge work | AutomationBench | 51.3% | 31.4% | +19.9 |
| Science | LABBench 2 | 88.8% | 68.6% | +20.2 |
| Long context | GraphWalks (256K to 1M) | 84.2% | 65.0% | +19.2 |
| Charts | Chartography | 71.6% | 46.2% | +25.4 |
| Video | LVBench | 91.7% | 79.7% | +12.0 |
| Coding | DeepSWE v1.1 | 77.9% | 67.4% | +10.5 |
| Security | CWE-bench v1 | 68.0% | 58.0% | +10.0 |
| Coding | FrontierSWE v2 | 55.0% | 56.3% | −1.3 |
| Coding | Terminal-bench 4.0 | 57.4% | 57.9% | −0.5 |
Fable 5.1's two wins are both demanding software-engineering tests, by 1.3 and 0.5 points: effectively a draw.
Among coding tests, the clear gap is DeepSWE v1.1 (long-horizon software engineering), more than 10 points. On Vibe Code Bench (building apps from instructions), Argon scores 91.9% and Fable 5.1 90.3%, so everyday coding shows little difference.
Does it also beat GPT-6 Astra and Claude Opus 5.5?
Of the table's 19 rows (GraphWalks is split into two context lengths), Argon leads outright on 13 and ties on one. Astra and Opus 5.5 each keep the lead in their strongest areas.
| Rows led | Model | Main benchmarks |
|---|---|---|
| 13 (outright) | Gemini 4 Argon | Vals Index, AutomationBench, DeepSWE v1.1, LABBench 2, Agent's Last Exam, and more |
| 3 (outright) | GPT-6 Astra | FrontierSWE v2 (65.5%), Terminal-Bench Science 0.1 (68.1%), OSWorld-2.0 (72.6%) |
| 2 (outright) | Claude Opus 5.5 | Terminal-bench 4.0 (66.4%), PostTrainBench (49.3%) |
| 1 (tie) | Argon and Astra | CWE-bench v1 (68.0%) |
Fable 5.1 does not lead any row. Within Anthropic's lineup, Opus 5.5, released on September 22, is stronger on agentic tests such as Terminal-bench 4.0, at an API price 60% lower than Fable 5.1.
One caveat about the table. When Anthropic launched Opus 5.5, it reported 55.8% for Fable 5.1 on Terminal-Bench 4.0; Google's table shows 57.9%. According to Google's methodology, the non-Gemini scores for this test come from the official public leaderboard, using each model's highest thinking setting. The choice of setting can move a score by a point or two, and the 0.5-point gap between Argon and Fable 5.1 is smaller than that.
Google compiled this table. According to its methodology document, results for non-Gemini models are the providers' self-reported numbers unless otherwise noted, and some come from public leaderboards. Argon's numbers were measured by Google, and no third party has reproduced them yet. Conditions may not match across tests, so treat gaps of a few points as rough indicators.
How accurate was the pre-launch leak?
The "beats Astra and Fable" leak that circulated about two weeks before launch got the direction right but the numbers wrong. According to TechBriefly, an unidentified model appeared on the Arena comparison site on September 17 under the name "gemini-3.8-flash," and developers concluded it was a Gemini 4 Pro checkpoint with the internal codename Argon.
| Item | Leak | Official |
|---|---|---|
| DeepSWE v1.1 | 88% | 77.9% |
| Output limit | 256,000 tokens | 1 million tokens |
| Price (input / output, per million tokens) | $2.25 / $11.25 | $2 / $10 (introductory) |
| Name | Gemini 4 Pro | Gemini 4 Argon |
The codename and the claim of beating Astra and Fable held up. The DeepSWE score, however, was more than 10 points above the official figure, and the output limit was a quarter of the real one.
We cannot tell whether the Arena model matched the final release; it may have been an earlier checkpoint. Either way, anyone who planned a migration or a budget around the leaked numbers would have gotten it wrong.
How should teams and enterprises approach Argon?
There is no need to switch today, but it is worth preparing a comparison now. Argon is not yet open to developers, so you cannot check whether the official numbers hold for your own workloads until access widens.
What to prepare depends on your work:
- Heavy legal, finance, or business-automation workloads: This is where Argon leads by the most. Pick about ten representative tasks you currently run on Fable 5.1 or Astra, and fix their inputs and expected outputs so you can compare under identical conditions on day one
- Teams that use coding agents daily: Opus 5.5 and Astra still lead in terminal work and frontier coding, and Fable 5.1 is roughly even with Argon there. Rather than a full switch, start with tasks that benefit from the 1-million-token output limit, such as long code migrations and large refactors
- Cost-sensitive teams: Introductory pricing is one-fifth of Fable 5.1 and Astra, but Google has not said how long it lasts, and afterwards Argon costs the same as Opus 5.5. Avoid designs that only work at the introductory price
September alone brought GPT-6 Astra (Sept. 3), Claude Opus 5.5 (Sept. 22), GPT-6.1 Sol (Sept. 29), and Gemini 4 Argon (Sept. 30). Building so that you can swap models later holds up better against this churn than depending deeply on one vendor.
FAQ
Q. When can I use Gemini 4 Argon?
A. As of October 1, 2026, only cyber defenders in Google's Fairwind Program and some trusted testers have access. Google says it will expand to developers, enterprises, and consumers "as soon as possible," starting with paid API customers and Google AI Ultra subscribers.
Q. Is Gemini 4 Argon better than Claude Fable 5.1?
A. In Google's table, Argon beats Fable 5.1 on 15 of the 17 comparisons where both have scores. Fable 5.1 wins FrontierSWE v2 and Terminal-bench 4.0, by 1.3 and 0.5 points.
Q. How much does Gemini 4 Argon cost?
A. During the introductory period, $2 per million input tokens and $10 per million output tokens, with cached input 95% off. Afterwards, $4 and $20. Google has not announced how long the introductory period lasts.
Q. Was the pre-launch leak accurate?
A. The codename Argon and the claim of beating Astra and Fable were right, but the DeepSWE score (88% in the leak, 77.9% officially) and the output limit were not.
Summary
In Google's official table, Gemini 4 Argon beats Claude Fable 5.1 on 15 of 17 comparisons and leads 13 of 19 rows outright. At one-fifth the introductory price of Fable 5.1 and Astra, it leads by wide margins in business automation, science, long context, and chart understanding. Opus 5.5 and Astra still lead in terminal work and frontier coding, where Fable 5.1 is roughly even with Argon, and most developers cannot use Argon yet.
At ZenChAIne, we use both Claude and GPT models in our daily development work. Once Argon becomes broadly available, we will share how it compares on real development tasks.
References
- Introducing Gemini 4 Argon - Sundar Pichai (X)
- Gemini 4 Argon: our next era of frontier intelligence - Google
- Gemini 4 Argon evaluation methodology - Google DeepMind
- Google unveils Gemini 4 Argon, retaking benchmark lead over OpenAI and Anthropic — but in limited release - VentureBeat
- Google announces Gemini 4 Argon as its new frontier model - 9to5Google
- Anthropic releases Claude Opus 5.5, beating Fable 5.1 on key agentic benchmarks at 60% cheaper API price - VentureBeat
- Leaked Gemini 4 Pro benchmarks show it beating GPT-6 Astra and Claude - TechBriefly
