Gemini 4 Argon: our next era of frontier intelligence
Google introduces Gemini 4 Argon, a frontier AI model designed for complex workflows in software engineering, enterprise tasks, and cybersecurity, with phased access and new pricing tiers.
Google has unveiled Gemini 4 Argon, a frontier AI model engineered to handle intricate, long-horizon tasks across software engineering, enterprise knowledge work such as legal and finance, and cybersecurity defense. The model is currently being distributed to trusted cyber defenders through the Fairwind Program as part of a controlled rollout. Argon is designed to sustain deep reasoning over extended workflows, enabling professionals to address complex challenges with greater efficiency and depth.
Gemini 4 Argon will launch at an introductory rate of $2 per million input tokens and $10 per million output tokens, with cached input tokens discounted by 95%. The model’s output token limit has been expanded to 1 million tokens, up from 64,000, allowing for deeper, multi-step reasoning in a single interaction. Google engineers are already leveraging Argon for tasks ranging from debugging to large-scale code migrations, citing significant improvements in productivity and problem-solving capabilities.
Argon has achieved state-of-the-art performance on specialized benchmarks, including DeepSWE v1.1 (77.9%) for software engineering, the Vals Index for economic impact across finance, legal, and tax work, and AutomationBench (51.3%) for end-to-end business automation. It also excels in visual understanding tasks, such as long video analysis on LVBench (91.7%), and cybersecurity defense, where it autonomously identifies and patches critical vulnerabilities.
To ensure safe deployment, Google is strengthening safeguards against misuse, prompt injection attacks, and misalignment while hardening sandboxed environments for secure testing. The model will initially be available to paid API customers and Google AI Ultra subscribers, with pricing increasing to $4 per million input tokens and $20 per million output tokens after the introductory period. Feedback from early testers will guide further refinements before broader release.