A launch with a narrow first audience

Google has announced Gemini 4 Argon, but the most important availability detail is what it has not yet done: the model is not a general release for every Gemini user. In its 30 September announcement, Google said Argon is rolling out first to trusted cyber defenders through its Fairwind programme. Developers, enterprises and consumers are slated for a later phase after the company gathers feedback and strengthens safeguards. For buyers comparing models today, the distinction between an announced capability and an accessible product is essential.

Google positions Argon as a frontier system for demanding, long-running work across software engineering, finance, legal tasks and cyber defence. It also says the model is being used internally by thousands of employees. These are vendor claims and early deployment examples rather than evidence that the wider market can already reproduce the outcomes. A sensible reading is that this is a significant model announcement with a deliberately limited initial distribution.

The unusual scale of its working room

Argon is designed for deep reasoning over complex workflows. Google says it has expanded the model’s output-token limit to one million tokens, up from a previous 64,000-token ceiling. That figure describes how much the model can generate in a trajectory, not a promise that every task needs or should consume that much output. Long trajectories can help with multi-file engineering or a sequence of research and verification steps, but they also increase the need to monitor costs, intermediate decisions and the quality of final work.

The announcement cites internal projects as examples. In one, Argon agents examined data-centre profiling telemetry and identified memory optimisations that Google says freed more than 300 TiB after rollout. Another involved moving C and C++ codebases towards Rust, including work on the Fuchsia Zircon kernel. Google stresses that critical rewrites undergo automated and manual auditing before production. That qualification matters: a model’s ability to propose a large change should not be confused with approval to deploy it without human review.

Benchmarks and business use cases

Google reports a 77.9 per cent result on DeepSWE v1.1 for long-horizon software engineering, a 51.3 per cent score on AutomationBench, and a 91.7 per cent score on LVBench for long-video understanding. It also claims leading results on evaluations covering finance, legal work and cyber vulnerability repair. Benchmark scores are useful signals about particular tests; they are not a substitute for evaluating an application with its own tools, documents, risk tolerance and operating costs.

For an enterprise team, the stronger practical question is whether Argon can finish a complete job reliably. A legal research assistant must cite the right authority, a finance workflow must preserve an audit trail, and a coding agent must pass tests and survive review. A model that can continue for longer may be particularly helpful on those multi-stage tasks, but it can also continue down an incorrect path for longer. Google’s examples illustrate the potential; production adoption will need task-level measurement.

Why cyber defence comes first

The first external cohort is focused on cyber defence. Google says Argon can find, validate and patch critical vulnerabilities, and reports that partner Wiz used it to identify a serious exposure in healthcare software that earlier frontier models missed. It cites a 68 per cent result on CWE-bench v1 for remediation. The company says selected trusted defenders and internal teams will receive access without the cyber guardrails applied to general users, making the control of that cohort especially consequential.

This choice reflects the dual-use nature of advanced security capability. A system that can identify flaws for defenders might also help an attacker if distributed without constraints. Google describes work on harmful-request refusals, prompt-injection robustness, monitoring for misalignment and hardened sandboxes. It says it is taking part in a United States government process for pre-release model access. None of these measures should be read as a guarantee against misuse; they explain why the broader release is staged.

Pricing is disclosed, access is still gated

Google lists introductory API pricing of US$2 per million input tokens and US$10 per million output tokens, with cached input tokens at a 95 per cent discount to the input rate. Its footnote says the later price will be US$4 and US$20 respectively after the introductory period. A team modelling cost should use the appropriate future price as well as the launch offer, particularly for tasks that generate very large outputs.

The company says paid API customers and Google AI Ultra subscribers are expected among the first broader groups after trusted testers. It has not provided a firm date for that expansion in the announcement. For now, the operational takeaway is to track access and safeguards separately from capability claims. Argon may become a major option for complex agentic work, but the model’s current release is a controlled trial of both performance and safety.