Gemini 4 Argon, explained: Google's new frontier model, and why I could not call it yet
Google announced Gemini 4 Argon on 30 September 2026, with a 1 million token output limit and a price of $2 in and $10 out per million tokens. It is only open to a small group of cyber defenders for now, so I checked our own API key, got a 404, and turned that into a lesson on choosing models in production.
What happened when our code asked for Argon on 4 October 2026: a quick 404, then a fallback model that answered. Section 5 has the full run.
In short
1Gemini 4 Argon is Google's new most capable model, announced on 30 September 2026 for long, many-step work in coding, finance, law and cyber defense.
2Outside Google, only an initial group of partners in Google's Fairwind Program can use it; Google says thousands of its own staff already do. Google says paid API customers and Google AI Ultra subscribers come next, with no date given.
3Google lists an introductory price of $2 per million input tokens and $10 per million output tokens, rising to $4 and $20 later. The output limit grows to 1 million tokens, from 64K.
4Our API key sees 61 models and none is Argon. Calling it returned HTTP 404, so my code fell back to Gemini 3.8 Flash, which answered all three test prompts correctly in 1.8 to 2.3 seconds.
5Every benchmark score here is Google's own. Nobody outside the early group can test Argon yet, including us.
Announced
30 September 2026
From
Google DeepMind
Who can use it
Fairwind partners and Google's own teams
Price at launch
$2 in, $10 out per million tokens
This page is the free part.
The course goes deeper on choosing and routing models in production
₹999 in India/$49 everywhere elseonce, for the whole course
The AI Engineering course covers choosing and routing models in production across a run of lessons, not one page. These 4 alone are about 315 minutes of step-by-step reading, every one with code you run in the browser, all with a quiz.
HTTP 429 is the first failure an LLM app hits at scale. Beat rate limits and provider outages with token buckets, backoff, degraded modes, and failover.
How to measure LLM output quality when there is no single correct answer: golden sets, LLM-as-judge, RAG metrics, online A/B tests, and regression testing every prompt change.
“Safely releasing frontier capabilities at this level requires a phased approach.”
Google, Gemini 4 Argon launch post
On 30 September 2026 Google published "Gemini 4 Argon: our next era of frontier intelligence", written by Koray Kavukcuoglu, a senior vice president at Google DeepMind. A frontier model is a company's most capable model, the one at the edge of what it can build.
Google says Argon is "Built to sustain deep reasoning across complex, long-horizon workflows". A long-horizon workflow is a task with many steps that can run for hours, such as moving a large codebase from one language to another, where the model must keep track of what it already did.
It is not open to everyone. The post says Argon is "rolling out to a set of trusted cyber defenders through our Fairwind Program". Fairwind is a Google DeepMind program that gives early access to governments, operators of critical services such as healthcare, telecoms and energy networks, and core technology platforms, to give them "a critical head start against cyber threats".
Google calls this a phased approach. It says it is taking part in the U.S. government's voluntary process for pre-release model access, and will widen access step by step, "starting with paid API customers and Google AI Ultra subscribers". The post gives no date for that.
2
What is different about Argon
The clearest change is the output limit. A token is a piece of text the model reads or writes, often part of a word. The output limit is the most tokens the model may write in one reply.
Google is raising that limit to 1 million tokens, "up from the previous 64K tokens". Our own API listing agrees on the old number: Gemini 3.8 Flash and Gemini 3.1 Pro Preview both report an output limit of 65,536 tokens, which is 64K.
Why does that matter? With a 64K limit, a model doing a very long job has to stop, hand over and be called again. Google's argument is that with room for hundreds of thousands of tokens in one reply, the model can think longer and finish hard problems "in one go". The flip side is cost: output tokens are the expensive ones, so a reply that long is a big bill.
Google also describes Argon at work inside Google. It says a team of Argon agents found memory savings across its data centers, "freeing up over 300 TiB of memory" (a TiB, or tebibyte, is about 1.1 trillion bytes), and that Argon agents took an existing Rust version of a video decoder and replaced 32,000 lines of SIMD code in it (code that does the same operation on many values at once), so it "runs 2.7x faster than the Rust port". Rust is a programming language designed to avoid memory bugs. These are Google's own accounts and we cannot check them.
One reply can now be about 15 times longer
Both bars use one scale. 65,536 is what our own API key reports for today's Gemini models; 1 million is Google's figure for Argon. Thinking tokens are billed as output, so long replies cost more.
3
How good is it? Google's numbers, not ours
A benchmark is a fixed set of test tasks with a score, so models can be compared. Every number below is from Google's launch post. We could not run Argon, so none of them is ours.
AutomationBench (Zapier): end-to-end business tasks
51.3%, ranked first
LVBench: understanding long videos
91.7%, which Google calls state of the art
CWE-bench v1: fixing security weaknesses in code
68%, tied for first
Vals Index: finance, coding, legal and tax work
described as the leading model, no score in the text
Gray Swan Indirect Prompt Injection benchmark
described as leading, no score in the text
Google's own figures from its 30 September 2026 launch post. Not reproduced by us or, as far as we found, by anyone independent.
Two cautions. A company picks which benchmarks to publish, so a launch post shows the strengths. And a high score on someone else's tasks does not mean a high score on yours, which is why the last section of this page is about testing on your own work.
4
Who can use Argon, and what it will cost
Here is what Google's launch post and Fairwind page say, read on 4 October 2026.
Plans and limits
Fairwind Program partners
access now
Paid API customers and Google AI Ultra
next, no date given
Everyone else, Gemini app and free API
not yet
Input, introductory price
$2 per million tokens
Output, introductory price
$10 per million tokens
Cached input
95% off the input price
Input and output after the introduction
$4 and $20 per million tokens
API model name
not published
* * * * *
Cached input means text you send again and again, such as a long instruction block at the start of every request, which the service stores and charges less for. 95% off $2 works out to $0.10 per million tokens; that is my arithmetic, not a number Google printed.
Google does not say how long the introductory period lasts. If you plan a budget around Argon, use the later price of $4 and $20, because that is what you will pay once the offer ends.
Google also says that for trusted defenders and its own teams it will release Argon "without cyber guardrails", so they can use its full ability to find and fix security holes.
5
My hands-on: asking for Argon, and falling back
Our site has its own Gemini API key. An API key is the secret that lets code call a model over the internet. First I asked the API which models this key can see. The answer was 61 models, and none has "argon" in its name.
Google has not published an API name for Argon. So I tried the three names a developer would most likely guess: gemini-4-argon, gemini-4-argon-preview and gemini-argon. All three came back HTTP 404 NOT_FOUND, the web's way of saying "no such thing here".
A 404 is exactly what a production app hits on launch day, when a new model is announced but not yet open to its key. So I wrote a small router: it asks for Argon first and, on a 404, a rate limit or a server error, asks the next model on the list. The fallback is Gemini 3.8 Flash, a current general-purpose model this key can see.
I gave it three small jobs. A reasoning question: the average response time of a cache with a 90% hit rate. A coding task: a palindrome check. And a long-context test: 400 made-up log lines, about 17,000 tokens, with one failed request hidden at line 288.
All three answers were right. The arithmetic gave 6.8 ms. The function passed five test cases I ran on it, including an empty string and "A man, a plan, a canal: Panama". And it found the hidden failure, request 100287 with status 503.
The cost of guessing wrong shows in the timings. Every call spent 0.58 to 0.64 seconds on the 404 before the real answer started. A real router should remember that a model is missing, for a few minutes at least, instead of paying that delay on every request.
What we measured: 4 October 2026, from India, with our own Gemini API key: 3 of 3 calls to gemini-4-argon returned HTTP 404 in 0.58 to 0.64 s. The fallback, gemini-3.8-flash, answered all 3 correctly in 1.80 to 2.26 s. The long-context prompt was 17,235 input tokens. A first run four minutes earlier gave the same three answers, with 404s in 0.57 to 0.75 s and answers in 1.61 to 2.44 s.
$ python3 list_models.py
61 models visible to this key
with 'argon' in the name: []
models/gemini-3.1-pro-preview input limit 1048576 output limit 65536
models/gemini-3.8-flash input limit 1048576 output limit 65536
# router.py: try the first-choice model, fall back on 404/429/5xx, record latency and token usage.
import json, os, time, urllib.request, urllib.error
KEY = os.environ["GEMINI_API_KEY"]
CHAIN = ["gemini-4-argon", "gemini-3.8-flash"] # first choice, then fallback
URL = "https://generativelanguage.googleapis.com/v1beta/models/{}:generateContent"
def call(model, prompt):
body = json.dumps({"contents": [{"parts": [{"text": prompt}]}]}).encode()
req = urllib.request.Request(URL.format(model), data=body,
headers={"x-goog-api-key": KEY, "Content-Type": "application/json"})
t0 = time.perf_counter()
try:
with urllib.request.urlopen(req, timeout=120) as r:
return r.status, json.load(r), time.perf_counter() - t0
except urllib.error.HTTPError as e:
return e.code, json.load(e), time.perf_counter() - t0
def ask(prompt):
for model in CHAIN:
status, data, secs = call(model, prompt)
print(f" {model}: HTTP {status} in {secs:.2f} s")
if status == 200:
u = data["usageMetadata"]
text = "".join(p.get("text", "") for p in data["candidates"][0]["content"]["parts"])
print(f" answered by {data.get('modelVersion')} tokens in {u.get('promptTokenCount')}"
f" thinking {u.get('thoughtsTokenCount', 0)} out {u.get('candidatesTokenCount')}")
return text
if status not in (404, 429, 500, 503):
raise SystemExit(f"stop: {status} {data['error']['message']}")
raise SystemExit("every model in the chain failed")
PROMPTS = json.load(open("prompts.json"))
for name, prompt in PROMPTS.items():
print(f"[{name}] {len(prompt):,} characters")
print(" reply: " + ask(prompt).strip().replace("\n", "\n "))
$ python3 make_prompts.py # writes prompts.json: two short prompts and 400 log lines with one planted failure
$ python3 router.py
[reasoning] 164 characters
gemini-4-argon: HTTP 404 in 0.64 s
gemini-3.8-flash: HTTP 200 in 2.26 s
answered by gemini-3.8-flash tokens in 46 thinking 338 out 48
reply: (0.90 × 2 ms) + (0.10 × 50 ms) = 1.8 ms + 5.0 ms = 6.8 ms
**Answer:** 6.8 ms
[coding] 145 characters
gemini-4-argon: HTTP 404 in 0.58 s
gemini-3.8-flash: HTTP 200 in 1.80 s
answered by gemini-3.8-flash tokens in 32 thinking 205 out 40
reply: ```python
def is_palindrome(s):
cleaned = [c.lower() for c in s if c.isalnum()]
return cleaned == cleaned[::-1]
```
[long-context] 28,128 characters
gemini-4-argon: HTTP 404 in 0.63 s
gemini-3.8-flash: HTTP 200 in 2.12 s
answered by gemini-3.8-flash tokens in 17235 thinking 176 out 23
reply: Request ID 100287 failed with status 503 and error `upstream_timeout`.
Real output from my Mac, with three changes for reading: router.py's source is shown inline above its output, the $ python3 make_prompts.py line and the comments after # are added, and the date lines are removed. The key is read from an environment variable and never printed. The reply text is exactly what the model returned, including its Markdown (plain-text formatting marks such as ** for bold).
Every request paid about 0.6 seconds to learn Argon was not there
Measured in my second run. The dashed part is pure waste once you know the model is missing, which is why routers cache a failure for a while instead of retrying it every time.
6
What these calls cost, and the same tokens at Argon's price
Gemini 3.8 Flash costs $0.75 per million input tokens and $3.75 per million output tokens until 31 December 2026, then $1.50 and $7.50. Google's pricing page says the output price includes thinking tokens, the hidden reasoning a model writes before it answers.
The table works out what my three calls cost on Flash, from the token counts the API returned. The last column puts the same token counts through Argon's introductory price. That is only a price comparison: Argon would think for a different number of tokens, so its real bill would differ.
Prompt
Tokens in
Thinking + out
On 3.8 Flash
Same tokens at Argon's $2 / $10
Reasoning
46
338 + 48
$0.0015
$0.0040
Coding
32
205 + 40
$0.0009
$0.0025
Long context, 400 log lines
17,235
176 + 23
$0.0137
$0.0365
Token counts from the usage numbers the API returned with each answer in my second run. Prices from Google's Gemini API pricing page and the Argon launch post, 4 October 2026. The arithmetic is mine.
Output price per million tokens, from each provider's own page
Two of these prices are introductory and will double. Output is where long reasoning spends money, so it is the number to compare first.
7
How to choose a model in production
A launch like Argon's is a good moment to set up model choice properly, because the newest model is often not one you can call yet. These are the questions I would answer, in order.
1
Can I call it at all?
Check the provider's model list from your own key, as I did, not the launch post. Announced is not the same as available, and a model can be open to some keys and not others.
2
Is it good at my task?
Keep a small set of real examples from your product with known right answers, and score each candidate model on them. This is called an eval. Public benchmarks are a starting point, not a decision.
3
How fast is it?
Latency is the time from request to answer. Measure it from where your users are, and look at the slow end, not only the average. My Flash calls took 1.8 to 2.3 seconds from India for very short answers; a model that thinks longer will take longer.
4
What will it cost per request?
Multiply your typical tokens in and out by the price per million, and include thinking tokens, which are billed as output. Budget at the price after any introductory offer ends.
5
How much can it read and write?
The context window is how much text the model can read in one request; the output limit is how much it can write. My log test used about 17,000 tokens of a 1,048,576 token window. Pick a limit that fits your largest real input with room to spare.
6
What happens when it fails?
Keep an ordered list of models, try the next one on a 404, a rate limit (HTTP 429, too many requests) or a server error, and remember a failing model for a while so you stop paying for the same error. This pattern is called a circuit breaker. Log which model answered, so you can see how often the fallback ran.
7
Will the answers still be good on the fallback?
A fallback is only safe if it passed the same eval. Otherwise you swap an error the user can see for a wrong answer they cannot.
If you want to build that list-of-models setup properly, my lesson "Model Routing and Cascades: Running Three Models Without Chaos" walks through routing, cascades and fallbacks step by step, and "Rate Limits and Provider Failure: The Number One Production Error" covers the errors that trigger them.
8
What I could and could not check
Everything about Argon itself comes from Google's launch post and the Fairwind Program page, read on 4 October 2026 and linked below. I found no Argon model card and no Argon entry on Google's Gemini API models or pricing pages that day.
I could not run Argon. The three model names I tried are guesses, because Google has not published one, so the 404s show that this key cannot reach Argon under those names, not that no name exists.
Everything in the hands-on section is from my own runs. One machine and three prompts are a demonstration, not a benchmark.
Questions
What is Gemini 4 Argon?
Google's newest frontier AI model, announced on 30 September 2026, built for long, many-step tasks in software engineering, finance, legal work and cyber defense. It can write up to 1 million tokens in one reply.
Can I use Gemini 4 Argon now?
Probably not. On 4 October 2026, outside Google, it was only open to an initial group of Fairwind Program partners, who work on cyber defense. Google says paid API customers and Google AI Ultra subscribers come next, without a date. Our own API key could not see it.
How much does Gemini 4 Argon cost?
Google lists an introductory price of $2 per million input tokens and $10 per million output tokens, with cached input 95% cheaper. After the introductory period it becomes $4 and $20. Google has not said when the introductory period ends.
What is the API model name for Gemini 4 Argon?
Google has not published one. gemini-4-argon, gemini-4-argon-preview and gemini-argon all returned HTTP 404 on our key on 4 October 2026.
Is Gemini 4 Argon better than GPT-6 Astra or Claude Opus 5.5?
Google's launch post reports leading scores on several benchmarks, such as 77.9% on DeepSWE v1.1. These are Google's own numbers and have not been reproduced independently, and no outside developer can test Argon yet.
What does the 1 million token output limit mean?
One reply can be up to 1 million tokens long, against 64K before. That lets the model work through long problems in one go, but output tokens are the expensive ones, so very long replies cost a lot.
Sources
What this explainer is based on, so you can check it.
This explainer sits on top of our AI Engineering course: 204 lessons on RAG, evals, agents, serving, security and MLOps, many built around a real experiment. 10 lessons are free to read, with no card needed.
course 2
AI Engineering
Take models from notebook to production, with labs on real models.