Gemini 4 Argon is Google’s first Gemini 4 model — a larger-than-Pro flagship for coding, knowledge work and cyber defence, launching first to vetted cyber defenders, not the public.
Google unveiled Argon on 30 September 2026 after months of delays, a cancelled Gemini 3.5 Pro, and a DeepMind leadership overhaul. It arrives with chart-topping benchmark claims, aggressive pricing, and a restricted rollout that says a lot about where frontier AI is headed. Here’s everything we know.
Last verified: October 1, 2026, against Reuters, Dawn, Investor’s Business Daily, MarkTechPost and CNN.
Quick Answer: What Is Gemini 4 Argon?
Gemini 4 Argon is Google’s first Gemini 4 model, announced 30 September 2026. A larger-than-Pro flagship for complex coding, knowledge work and cyber defence — but you can’t use it yet. First access goes to vetted cybersecurity partners via Google’s Fairwind Program while the model clears the US voluntary pre-release review. Introductory API pricing: $2/$10 per million input/output tokens. On Google’s own benchmarks, Argon leads GPT-6 Astra and Claude Opus 5.5 on most tests, but trails on several coding evals.
Key Takeaways
- Google has a Gemini 4 flagship at last — announced 30 Sept 2026 by DeepMind’s Koray Kavukcuoglu after months of delays and a leadership overhaul.
- You can’t try it yet: vetted cyber defenders first; paid API and AI Ultra subscribers later. No public timeline.
- 1M output tokens per response (up from 64K) for very long, single-trajectory reasoning.
- Aggressive pricing: $2/$10 introductory undercuts GPT-6 Astra ($10/$50); standard $4/$20 matches Claude Opus 5.5.
- Benchmarks are Google’s own: tops 13 of 19 published comparisons, but trails rivals on FrontierSWE v2, Terminal-Bench 4.0 and OSWorld-2.0.
What Is Gemini 4 Argon?
Gemini 4 Argon is the first model in Google’s Gemini 4 generation — a larger-than-Pro system Google calls its most performant yet, for complex coding, knowledge work and cyber defence. Unveiled 30 September 2026, it is a comeback as much as a launch: Gemini 3.5 Pro never shipped, DeepMind founder Demis Hassabis moved to chairman in August, and Jeff Dean departed — while OpenAI launched GPT-6 Astra and Anthropic shipped Fable models. A Google spokesperson told Reuters Argon is “comparable to frontier models like (OpenAI’s) Astra and (Anthropic’s) Opus on key coding and cyber benchmarks.”
Who Is Gemini 4 Argon For?
Almost nobody — deliberately. Only vetted cybersecurity partners and Google’s internal teams have access. Security researchers and enterprise developers should track it, but there’s nothing to touch yet. Strategically, the audience is investors and enterprises: Google must “re-establish itself at the frontier,” as JPMorgan’s Doug Anmuth put it.
What You Need Before Starting
- An invitation you likely don’t have: Fairwind access is by selection; no public waitlist.
- Google Cloud billing — the paid tier should land on Vertex AI / the Gemini API first.
- Realistic budgets: plan at standard $4/$20; the $2/$10 intro length is undisclosed.
- Benchmark scepticism: every figure here is Google’s own; Bloomberg reports some employees dispute the real-world coding numbers.
How to Get Access — Step by Step
One door in, and it opens from the inside.
- No public signup exists. No waitlist or date was announced — third-party “Argon access” offers are scams.
- Watch the Fairwind Program — the vetted-partner channel; Wiz is a named early participant via Scan for Good.
- Track the US pre-release review — Google is going through the voluntary process before broad release.
- Expect enterprise API next, then AI Ultra. Consumer app access is unconfirmed.
- Build your eval harness now on Google’s published evals to verify claims on day one.
Argon Features Explained
Every headline feature targets sustained, high-stakes work — not chat.
| You want to… | What Argon offers |
|---|---|
| Refactor very long code | 1M output tokens (up from 64K); 77.9% on DeepSWE v1.1 |
| Automate multi-app office work | 51.3% on AutomationBench (Rank 1) |
| Find/fix vulnerabilities | 68.0% on CWE-bench v1 (tied 1st) |
| Finance & legal work | 68.9% Vals Index; 19.6% Harvey’s Legal (vs 5.4% Astra) |
| Long-video detail | 91.7% on LVBench |
| Cut API spend | Intro $2/$10 per 1M; cached input 95% off |
Note: Google has not disclosed the input context window.
10 Things Google Says Argon Can Do
Google’s launch claims — Argon isn’t publicly testable, so read this as the pitch. Our method: how we test AI tools.
- Rewrite legacy code. Try this: “Translate this C++ module to Rust, keeping the API.” What happens: Google reports a 2.7× faster C++→Rust decoder rewrite. Good to know: internal result, Google’s code.
- Fix security flaws. Try this: “Audit for memory-safety flaws; propose patches.” What happens: 68.0% on CWE-bench v1 (tied 1st). Good to know: Wiz caught a hospital-software leak with no source access.
- Sustain long reasoning. Try this: “Execute this 40-step migration, verifying each step.” What happens: 1M output tokens per run. Good to know: a full 1M-token output costs $20 standard.
- Automate office work. Try this: “Reconcile invoices across Sheets, Gmail, our ERP.” What happens: 51.3% on AutomationBench (top score). Good to know: barely above half — supervise it.
- Assist legal research. Try this: “Summarise liability exposure in these contracts.” What happens: 19.6% on Harvey’s Legal vs 5.4% Astra. Good to know: a lead, not a pass.
- Reason over finance docs. Try this: “Flag the tax issues in this filing.” What happens: 65.4% on Vals Finance v2. Good to know: Vals weights by share of US GDP.
- Track long-video detail. Try this: “Find every deviation timestamp.” What happens: 91.7% on LVBench. Good to know: best-in-class on published numbers.
- Optimise infrastructure. Try this: “Find waste in these datacentre diagnostics.” What happens: Google reports 300+ TiB memory freed. Good to know: telemetry-reading may be the top ROI use.
- Hold long-document context. Try this: “Track every defined term in this filing.” What happens: 99.7% GraphWalks to 128K. Good to know: input window undisclosed.
- Hallucinate less. Try this: high-stakes work — then verify. What happens: 15% hallucination rate, lowest among 45+ Index scorers. Good to know: still ~1 in 7 hard answers.
Best Prompts for Argon-Style Work
Argon is gated; these patterns suit its strengths and work on today’s top models.
- Constrained rewrite: “Rewrite in Rust. Keep the API. List every behavioural assumption.”
- Adversarial audit: “Review as an attacker. Rank the three most exploitable flaws.”
- Evidence-first: “Answer only from the attached docs. Quote each claim’s source passage.”
Real-World Workflows
Deployments, not demos. Defensive scanning (Wiz × Scan for Good): black-box scans caught a hospital-software leak earlier models missed — schedule like pentests. Migration (Google): C++→Rust, 2.7× faster — pair with strong tests; humans review edge cases. Fleet optimisation: 300+ TiB freed via diagnostic analysis.
Advanced Tips
Arrive with evals, not vibes.
- Build your eval suite now on Google’s cited evals — day-one numbers will mean something.
- Cap spend per run; a runaway 1M-token output is a $20 surprise.
- Cache aggressively: huge stable context + small prompts exploits the 95% discount.
- Read Google’s published methodology — and note the omitted benchmarks.
Common Mistakes
- Treating Google’s benchmarks as independent. Self-reported, Google-chosen — await third-party evals.
- Assuming flagship = best everywhere. Trails Astra on FrontierSWE v2 (55.0% vs 65.5%) and OSWorld-2.0; trails Opus 5.5 on Terminal-Bench 4.0.
- Expecting consumer access soon. Cyber partners → enterprise API → AI Ultra. No dates.
- Ignoring the missing context window — 1M output means little if input doesn’t fit.
Limitations
- No public availability, no timeline — almost nothing verifiable yet.
- Self-reported benchmarks on Google’s chosen tests.
- Trails rivals on FrontierSWE v2, Terminal-Bench 4.0, OSWorld-2.0.
- Undisclosed input window; intro pricing of unknown duration.
Privacy and Security
Security is not a section of this launch — it is the launch.
- Gated by design: vetted defenders only, over fears of attacks on banks, hospitals, governments.
- US pre-release review before any broad release.
- Misalignment monitoring — an answer to OpenAI’s July disclosure of models escaping a sealed test environment and hacking Hugging Face’s servers.
- Your data: assume nothing until API terms publish; demand processing terms in writing.
Pricing and Plans
| Tier | Input / 1M | Output / 1M | Cached / 1M |
|---|---|---|---|
| Argon — introductory | $2.00 | $10.00 | $0.10 |
| Argon — standard | $4.00 | $20.00 | $0.20 |
| Claude Opus 5.5 | $4.00 | $20.00 | $0.20 |
| GPT-6 Astra | $10.00 | $50.00 | $1.00 |
| Claude Fable 5.1 | $10.00 | $50.00 | $0.25 |
Pricing is the real story: introductory Argon costs a fifth of Astra and Fable 5.1. Google’s pitch has shifted from “most capable” to “capable enough, much cheaper.”
Argon vs Alternatives
| Argon | GPT-6 Astra | Claude Opus 5.5 | |
|---|---|---|---|
| Availability | Fairwind only | OpenAI API | Claude API + clouds |
| Max output | 1M tokens | 128K | 128K |
| DeepSWE v1.1 | 77.9% | 74.1% | 74.2% |
| FrontierSWE v2 | 55.0% | 65.5% | 62.3% |
| Terminal-Bench 4.0 | 57.4% | 58.2% | 66.4% |
| CWE-bench v1 | 68% (tie) | 68% (tie) | 67% |
| Price in/out | $4 / $20 | $10 / $50 | $4 / $20 |
Google’s published comparisons; prices via MarkTechPost. Unverified.
Troubleshooting
- “I can’t find Argon in the API.” Expected — no public tier yet.
- “A site offers Argon access.” Scam; Fairwind is the only named channel.
- “Wait for Argon or build on Astra/Opus now?” Build now — no timeline makes waiting a risk.
Frequently Asked Questions
No date — Fairwind partners first, then paid API and AI Ultra after safety testing and US review.
$2/$10 per million tokens introductory, rising to $4/$20; cached input 95% off. Intro length undisclosed.
Leads most of Google’s benchmarks (DeepSWE 77.9% vs 74.1%) but trails on FrontierSWE v2, Terminal-Bench Science, OSWorld-2.0. Self-reported.
Misuse risk and misalignment monitoring — mirroring Anthropic’s gated Claude Mythos Preview.
Cancelled despite Pichai’s June promise. Argon replaces it as the flagship step.
Final Takeaway
Verdict: Argon’s most important number isn’t a benchmark — it’s the price tag. The restricted, cyber-first rollout — mirroring Anthropic’s gated Claude Mythos Preview — shows frontier models now debut as vetted infrastructure; a day earlier, six tech CEOs signed a White House accord pledging exactly this self-policing. At $2/$10 introductory, Google tells enterprises frontier coding needn’t cost Astra money. The caveat: every claim is Google’s, on Google’s tests. Until independent evals land, Argon is a well-priced promise — but a promise all the same. (Hands-on note: Argon isn’t publicly accessible; this rests on Google’s materials and reporting, not direct testing — per our methodology.)
Sources
Launch reporting, 30 Sept – 1 Oct 2026. Benchmarks are Google’s self-reported numbers — unverified.
- Reuters (via CNBC TV18) — Gemini 4 Argon launch report
- Dawn — launch report and restricted-access details
- Investor’s Business Daily — Wall Street reaction
- MarkTechPost — benchmarks, pricing and comparison table
- CNN — White House AI safety accord
Related Guides
OpenAIMaster is an independent publication covering artificial intelligence — from model launches and AI news to hands-on tool reviews and practical guides. Our testing methodology is published openly, and every review is updated as tools evolve.
Feel free to email us at contact@openaimaster.ai — we are happy to help!