
TLDR
Google's Gemini 4 Argon arrives priced at US$2 per million input tokens and US$10 per million output tokens. The output-token ceiling rises from 64,000 to 1 million, with early access reserved for cybersecurity partners through the Fairwinds program.
A bigger model, a longer context window
Google's Gemini 4 Argon landed on 30 September 2026 with one specification that signals an immediate architectural shift: the output-token limit jumps to 1 million tokens, up from a previous ceiling of 64,000.[1] That is roughly a 15-fold increase in how much the model can generate in a single call, which matters most to teams running long-form document analysis, extended code generation or multi-step legal review.
A Google spokesperson said Argon is larger than the company's previous top-tier Pro models and targets complex workloads, sitting alongside OpenAI's Astra and Anthropic's Opus on key coding and cybersecurity benchmarks.[2] Google DeepMind SVP Koray Kavukcuoglu confirmed the rollout. He said the model is going first to a set of trusted cyber defenders through the Fairwinds program and is designed for real-world software engineering, enterprise knowledge work and cyber defence.[1]
Fairwinds first, then wider API access
Access at launch runs through Google's Fairwinds program, which prioritises trusted cyber-defence partners before the model opens to general API customers and Google AI Ultra subscribers.[1] The phased rollout follows the same pattern Google used for earlier Gemini releases, slowing broad access while partners stress-test the model on live security workflows.
Google's introductory API pricing sits at US$2 per million input tokens and US$10 per million output tokens, with cached input tokens discounted 95 per cent off the standard input price.[1] That cache discount reshapes the economics for teams doing repeated reasoning over large, stable documents. A legal team running the same contract corpus through dozens of queries pays close to nothing on the input side after the first pass.
Benchmark scores and an internal dispute
On CWE-bench v1, which evaluates a model's ability to remediate security vulnerabilities, Argon tied for first place with a score of 68 per cent.[3] That result puts Argon alongside the current leaders in automated vulnerability patching, the subset of security work where model output most directly displaces human engineer hours.
Some Google staff told Bloomberg the model struggles with certain real-world coding tasks, a characterisation the company disputed. Google's release material describes frontier performance in complex workflows across real-world software engineering as a primary design goal.[1] The Fairwinds rollout period is effectively the live test of whether benchmark scores hold under production conditions.
What changes for developers, search teams and marketers
For developers calling the API directly, a 1-million-output-token ceiling means agentic coding pipelines that previously chained multiple calls and stitched outputs together can now close that loop inside a single request.[1] The per-call scope expands by an order of magnitude relative to earlier Gemini models.
Argon feeds into Google's AI Mode and ad creative workflows, so the underlying model driving AI-generated overviews and creative generation becomes substantially more capable at longer-form outputs.[1] Google has not confirmed a general availability date beyond the Fairwinds launch window.
KEY TAKEAWAYS
SOURCES & CITATIONS
FREQUENTLY ASKED QUESTIONS
What is the API price for Gemini 4 Argon?
Who gets access to Gemini 4 Argon first?
How does Argon perform on security benchmarks?
What is the output-token limit for Gemini 4 Argon?

Zara Kincaid writes about artificial intelligence and search. Her focus is what happens to businesses when the front page of the internet stops being a list of links and starts being an answer.




