Last updated: October 2026
What is Gemini 4 Argon, and can you use it yet?
Gemini 4 Argon is Google DeepMind's new frontier model, announced on September 30, 2026. Today it is limited to trusted cyber defenders through the Fairwind Program, with paid API customers and Google AI Ultra subscribers next and no date given, according to Google (2026).
This guide covers access, pricing, vendor-reported and independent benchmarks, and the security questions a CISO should ask before approving the model. Every figure is dated and linked to the page where it appears.
TL;DR: Key Takeaways
- Access is narrow. Gemini 4 Argon reached only Fairwind cyber defenders at launch, and DataCamp (2026) found no published API model ID on September 30.
- Price undercuts the rivals at launch. Google lists $2 per million input tokens and $10 per million output tokens, rising to $4 and $20, with a 95% discount on cached input (Google, 2026).
- Output is the headline spec. The output limit is 1 million tokens, up from 64,000, and Google has not stated an input context window (MarkTechPost, 2026).
- Benchmarks are mixed. Google reports 77.9% on DeepSWE v1.1 against 74.2% for Claude Opus 5.5, yet Opus 5.5 leads Terminal-Bench 4.0 at 66.4% versus 57.4%, and no third party had reproduced Google's table at launch (DataCamp, 2026).
- Independent boards disagree. Artificial Analysis scores Argon 52.6 against 57.6 for Opus 5.5, while Vals ranks it first of 41 at 68.9% (Trending Topics, 2026; Vals AI, 2026).
- Security claims need testing. Google says Argon leads Gray Swan's indirect prompt injection benchmark but published no score, so enterprises should run their own tests.
At a glance: Argon against Opus 5.5 and GPT-6 Astra
| Gemini 4 Argon | Claude Opus 5.5 | GPT-6 Astra | |
|---|---|---|---|
| Announced | Sept 30, 2026 | Sept 22, 2026 | Sept 3, 2026 |
| Public access | Fairwind cyber defenders only | Publicly available | Not verified here |
| Input / output price per 1M tokens | $2 / $10 intro, then $4 / $20 | $4 / $20 | $10 / $50 (VentureBeat) |
| Output limit | 1M tokens | 128K tokens | 128K tokens |
| DeepSWE v1.1 (vendor-reported) | 77.9% | 74.2% | 74.1% |
| Vals Index (vendor-reported) | 68.9% | 67.0% | 63.1% |
| Terminal-Bench 4.0 (vendor-reported) | 57.4% | 66.4% | 58.2% |
| Artificial Analysis Intelligence Index | 52.6 | 57.6 | 52.7 |
Sources: Google, DataCamp, VentureBeat, Trending Topics, MarkTechPost (output limits), NeuralTrust. Retrieved October 1, 2026.
What is Argon and who can use it today?
Gemini 4 Argon is a Google DeepMind frontier model announced on September 30, 2026 and released first to trusted cyber defenders through the Fairwind Program. Google says paid API customers and Google AI Ultra subscribers follow "as soon as possible", while the model also takes part in the U.S. government's voluntary pre-release review (Google, 2026).
Where Argon sits in the Gemini line
Google's last flagship series was Gemini 3 in November 2025, according to VentureBeat (2026). Fello AI (2026) adds that Argon is the first Gemini with a codename instead of the Pro, Flash and Flash-Lite tiers, and that Google has not said what other Gemini 4 models will be called.
Who has access right now
DataCamp (2026) reports that on launch day Argon was absent from OpenRouter, Vertex AI, Gemini CLI, Cursor and GitHub Copilot. Engineering teams cannot move production traffic to it yet. Security teams are the exception: Google says Wiz used Argon through its Scan for Good initiative and found a critical vulnerability in healthcare software that earlier frontier models had missed.
Argon pricing: what does it cost against Opus 5.5 and GPT-6 Astra?
During the introductory period Gemini 4 Argon costs $2 per million input tokens and $10 per million output tokens. Standard rates are $4 and $20, with a 95% discount on cached input. Google has not said how long the introductory price lasts. At standard rates Argon matches Claude Opus 5.5.
| Model | Input per 1M tokens | Output per 1M tokens | Cached input per 1M tokens | Cost of 1M input + 200K output |
|---|---|---|---|---|
| Gemini 4 Argon (introductory) | $2 | $10 | $0.10 | $4 |
| Gemini 4 Argon (standard) | $4 | $20 | $0.20 | $8 |
| Claude Opus 5.5 | $4 | $20 | $0.20 | $8 |
| GPT-6 Astra | $10 | $50 | not reported | $20 |
Sources: Fello AI (cached prices), VentureBeat (Opus 5.5 and Astra prices), NeuralTrust. The last column is our arithmetic from the listed rates.
Sticker price is only part of the bill. According to Trending Topics (2026), Artificial Analysis measured a cost per average task of $1.99 for Argon, $3.26 for GPT-6 Astra and $5.98 for Claude Opus 5.5. The same evaluation shows Argon used 110 million tokens to run the index against a median of 82 million, so verbosity eats into the saving. Artificial Analysis also notes that Argon's price is introductory. DataCamp (2026) adds that Google has not said how reasoning tokens are billed.
Gemini 4 Argon benchmarks: what Google reports
Google reports Gemini 4 Argon ahead on DeepSWE v1.1 at 77.9%, the Vals Index at 68.9%, AutomationBench at 51.3% and Harvey's Legal Agent benchmark, and behind on FrontierSWE v2 and Terminal-Bench 4.0. These are vendor-reported figures, and DataCamp (2026) notes that no third party had reproduced them at launch.
| Benchmark (vendor-reported) | Gemini 4 Argon | GPT-6 Astra | Claude Fable 5.1 | Claude Opus 5.5 |
|---|---|---|---|---|
| DeepSWE v1.1 | 77.9% | 74.1% | 67.4% | 74.2% |
| Vals Index | 68.9% | 63.1% | 65.8% | 67.0% |
| Harvey's Legal Agent | 19.6% | 5.4% | 6.7% | 3.8% |
| FrontierSWE v2 | 55.0% | 65.5% | 56.3% | 62.3% |
| Terminal-Bench 4.0 | 57.4% | 58.2% | 57.9% | 66.4% |
| CWE-bench v1 | 68.0% | 68.0% | 58.0% | 67.0% |
| GraphWalks (256K-1M) | 84.2% | 71.8% | 65.0% | 66.8% |
Source: DataCamp (2026), compiling Google's announcement. Retrieved October 1, 2026.
The pattern is specific. Argon leads on long-horizon software tasks, legal and finance work, long-context retrieval and business automation. MarkTechPost (2026) reports 51.3% on AutomationBench against 42.5% for Opus 5.5. It trails where agents work in a terminal or across a large codebase: a 10.5-point gap to GPT-6 Astra on FrontierSWE v2 and a 9-point gap to Opus 5.5 on Terminal-Bench 4.0. Google also reports 91.7% on LVBench for long video understanding.
How we compared
We did not run these benchmarks. We use vendor-reported numbers where they are the only source and label them as such. Where an independent board exists we show it beside the vendor figure. Prices come from the vendors or from a named outlet, and every table carries its retrieval date.
What independent evaluations say about Argon
Independent evaluations place Gemini 4 Argon at or near the top on some boards and behind Claude Opus 5.5 on others. Artificial Analysis scores it 52.6 against 57.6 for Opus 5.5, while Vals ranks it first of 41 models on its index at 68.9%. The honest summary is a strong, uneven model.
| Independent measure | Gemini 4 Argon | Claude Opus 5.5 | GPT-6 Astra |
|---|---|---|---|
| Artificial Analysis Intelligence Index | 52.6 | 57.6 | 52.7 |
| Artificial Analysis cost per average task | $1.99 | $5.98 | $3.26 |
| Terminal-Bench 4.0 (Artificial Analysis) | 57.1% | 59.6% | 59.1% |
| Humanity's Last Exam (Artificial Analysis) | 57.1% | 61.4% | 54.7% |
| Vals Index | 68.9% | 67.0% | not listed |
| Vals cost per test | $15.68 | $32.14 | not listed |
Sources: Trending Topics (2026) for Artificial Analysis, Vals AI (2026). Artificial Analysis tested only Argon's High setting. Retrieved October 1, 2026.
The Decoder (2026) reports a hallucination rate of 15% for Argon against 51% for GPT-6 Astra, but factual accuracy of 50% against 63%, and a first place on the Arena.ai Text Arena at 1,525 points. It concludes that Google closes the gap without taking a clear lead. Axios (2026) adds that Bloomberg reported some Google staff found performance lacking in internal testing, a claim Google disputed.
Context window and 1M-token output limit
Google states a 1 million token output limit for Gemini 4 Argon, up from 64,000 on earlier Gemini models, and does not state an input context window. Some outlets report a 1 million token input window, but we could not match that figure to a Google source, so treat it as unconfirmed.
The output figure is the real differentiator. MarkTechPost (2026) reports that Claude Opus 5.5, Claude Fable 5.1 and GPT-6 Astra each cap a single response at 128,000 tokens. Google cites internal migrations of C and C++ code to Rust at up to 800,000 lines, including the Fuchsia Zircon kernel, as the target use case.
A response of that size has a cost and a risk. At $10 per million output tokens a maximal answer costs up to $10 at introductory pricing and $20 at standard pricing. More important, an agent that writes 800,000 lines in one pass produces a diff no human can review line by line. The output limit raises the stakes of every guardrail around the agent.
Gemini 4 Argon security: what enterprises need to know
Google reports that Gemini 4 Argon leads Gray Swan's indirect prompt injection benchmark and ships with chain-of-thought and action monitors that can stop execution. It published no injection score, and DataCamp (2026) reports that the Fairwind cohort received the model without cyber safeguards. Model-level defenses help, but they do not replace controls around the agent.
What Google says it built in
Google lists refusal training for harmful requests, safeguards for cyber and CBRN misuse under its Frontier Safety Framework, internal and external red teaming, and activation monitoring to detect misuse. It adds automated red teaming and adversarial training against prompt injection, plus monitors that watch reasoning and actions and can halt a run. DataCamp (2026) notes Google kept monitoring findings out of training so the model does not learn to evade them.
Risks that remain for deployers
According to OWASP (2025), prompt injection is the first risk (LLM01) in its Top 10 for LLM applications, and indirect injection through documents, web pages and tool output is the form that reaches agents. A "leads the benchmark" claim is also hard to compare. NeuralTrust's Opus 5.5 analysis (2026) reports that Opus 5.5 ties Claude Fable 5.1 for the lowest injection success rate on the same Gray Swan benchmark, so several vendors can claim a lead without published scores to settle it.
Three enterprise exposures follow. Prompt injection can redirect an agent that has tool access. Data can leak through long outputs and tool calls. And an agent with write permissions can act on a mistaken or manipulated plan at machine speed. NeuralTrust's GPT-6 Astra briefing (2026) describes the same shift for OpenAI's model, the first to reach OpenAI's Critical cybersecurity threshold. Frontier cyber capability is now a feature of the whole market, not of one vendor.
How NeuralTrust addresses this
Control has to sit where the agent acts, whichever model sits behind it. Agent Gateway (TrustGate) enforces policy between agents, tools and any model provider, so swapping Claude for Gemini does not change the rules. Agent Runtime Security (TrustGuard) inspects prompts, tool calls and outputs as they happen. AI Red Teaming (TrustTest) lets you run your own injection and misuse tests against Argon before you approve it, instead of relying on a vendor's claim. Agent Posture Management (TrustLens) shows which agents and tools have access to what. NeuralTrust delivers these as the Runtime Security Mesh, and has four Gartner Hype Cycle 2026 recognitions in AI Runtime Defense.
Which should you choose?
Choose by what you can deploy today and what you can verify. Claude Opus 5.5 is available now and leads the independent index. Argon is a strong candidate for long, document-heavy work once it reaches general availability. Treat any switch as a test, not a swap.
| If you are... | Choose | Why |
|---|---|---|
| A security team that needs a model in production this month | Claude Opus 5.5 | Publicly available, highest Artificial Analysis score |
| A platform team running high-volume, cost-sensitive pipelines | Evaluate Argon at general availability | Lower introductory price, but verbose and price may change |
| A legal or finance team running research agents | Test Argon first | Leads Harvey's benchmark and the Vals Index |
| An engineering team running terminal-based coding agents | Claude Opus 5.5 | Leads Terminal-Bench 4.0 on both vendor and independent runs |
| A team planning very large code migrations | Watch Argon | 1M output limit, with review and rollback controls in place |
| A vetted cyber defender | Apply to the Fairwind Program | Only route to Argon today |
For model-by-model security detail, read our Claude Opus 5.5 enterprise security analysis and the GPT-6 Astra CISO guide. For the broader buying checklist, see the CISO guide to AI security.
Conclusion
Gemini 4 Argon is a credible frontier model with a price advantage at launch, a 1 million token output limit and strong results on legal, finance and automation tests. It is also unavailable to most teams, unproven outside Google's own benchmark table, and behind Claude Opus 5.5 on the independent index. Evaluate it when access opens, and put your own controls around it before it touches production data.
Secure Gemini and Multi-Model AI Agents in Production with NeuralTrust
Run Gemini, Claude and GPT behind one policy layer and test each model against your own threats before it goes live.
Related Comparisons
- Gemini vs. ChatGPT: Benchmark & Models Comparison 2026
- Claude Opus 5.5 vs. Gemini: Benchmark Comparison 2026
- Claude Opus 5.5 Enterprise Security: Safeguards & Gaps
FAQs about Gemini 4 Argon
1. What is Gemini 4 Argon?
Gemini 4 Argon is Google DeepMind's frontier model, announced on September 30, 2026. It is the first Gemini with a codename in place of the Pro and Flash tiers. Google positions it for coding, knowledge work and cyber defense, with a 1 million token output limit and a 95% discount on cached input (Google, 2026).
2. Is Gemini 4 Argon available to the public?
Not yet. At launch only trusted cyber defenders in the Fairwind Program had access. Google says paid API customers and Google AI Ultra subscribers come next without giving a date. DataCamp (2026) found no published API model ID and no listing in OpenRouter, Vertex AI or Cursor on September 30.
3. How much does Gemini 4 Argon cost?
Introductory pricing is $2 per million input tokens and $10 per million output tokens. Standard pricing is $4 and $20, and cached input carries a 95% discount. Google has not said how long the introductory rate lasts or how reasoning tokens are billed, so budget for the standard rate until it does (Google, 2026).
4. What is the Gemini 4 Argon context window?
Google states a 1 million token output limit, up from 64,000, but does not state an input context window in its announcement. Some outlets report 1 million tokens of input, which we could not trace to a Google source. Check the model documentation once the API opens before you design around a figure (MarkTechPost, 2026).
5. Is Gemini 4 Argon better than Claude Opus 5.5?
It depends on the measure. Google reports Argon ahead on DeepSWE v1.1 (77.9% against 74.2%) and the Vals Index, while Opus 5.5 leads Terminal-Bench 4.0 and the Artificial Analysis Intelligence Index, 57.6 against 52.6. Argon is cheaper at launch but not yet widely available (Trending Topics, 2026).
6. Is Gemini 4 Argon better than GPT-6 Astra?
On Google's table Argon beats GPT-6 Astra on DeepSWE v1.1, the Vals Index and Harvey's Legal Agent, and trails it on FrontierSWE v2 by 10.5 points. Artificial Analysis rates the two almost level, 52.6 against 52.7. Argon's introductory price is one fifth of Astra's, according to VentureBeat (2026).
7. Is Gemini 4 Argon safe for enterprise use?
Google reports leading results on Gray Swan's indirect prompt injection benchmark and describes reasoning and action monitors, but it has not published a score. No model is injection-proof. Enterprises should test Argon against their own threats and enforce policy at the gateway and runtime layers before granting it tool access (DataCamp, 2026).
8. What is the Fairwind Program?
The Fairwind Program is Google DeepMind's route for giving trusted cyber defenders early access to Gemini 4 Argon. DataCamp (2026) reports that the Fairwind cohort received the model without its cyber safeguards. Google's announcement links to the program page for details (Fairwind Program).
About the Author
Roger Howroyd is Head of Global SEO and AI at NeuralTrust, where he leads the company's search strategy across SEO, AEO, GEO, and LLM optimization. He specializes in AI-powered search, content strategy, and SEM. Connect on LinkedIn.
NeuralTrust is the leading platform for securing and scaling AI agents. Named a Pioneer in the Gartner Emerging Market Quadrant for AI Application Security 2026, recognized across four Gartner Hype Cycle reports in 2026, and featured in the Gartner Market Guide for Guardian Agents 2026, the Gartner Market Guide for AI Gateways 2025 and the KuppingerCole Leadership Compass for Generative AI Defense 2025. Headquartered in Barcelona with offices in London and New York. ISO 27001 certified.
Sources
- Google, Gemini 4 Argon: our next era of frontier intelligence, September 30, 2026.
- DataCamp, Gemini 4 Argon: Benchmarks, Pricing, and Access, September 30, 2026.
- MarkTechPost, Google DeepMind unveils Gemini 4 Argon with 1M output tokens, September 30, 2026.
- VentureBeat, Google unveils Gemini 4 Argon, retaking benchmark lead in limited release, 2026.
- Trending Topics, Gemini 4 matches GPT-6 Astra but trails Opus 5.5, 2026.
- Vals AI, Gemini 4 Argon benchmarks, cost and capabilities, September 30, 2026.
- The Decoder, Google Gemini 4 Argon closes the gap with OpenAI and Anthropic, 2026.
- Axios, Google unveils Gemini 4, September 30, 2026.
- Fello AI, Gemini 4 Argon: Benchmarks, Price and Who Gets It, 2026.
- OWASP, LLM01: Prompt Injection, 2025.
)
)
)
)
)
)
)