Headlines

Google unveils Gemini 4 Argon, its most powerful AI model yet; CEO Sundar Pichai says: We’re going to make it available as soon as we can and as…


Google unveils Gemini 4 Argon, its most powerful AI model yet; CEO Sundar Pichai says: We're going to make it available as soon as we can and as...
Google takes on GPT-6 Astra and Claude with Gemini 4 Argon.

Google has unveiled Gemini 4 Argon, its most capable AI model so far and the first in the Gemini 4 series. Announced on September 30, the new frontier model is built for long, multi-step work across software engineering, legal and finance tasks, and cybersecurity defence. It also ends a long wait, after the company quietly dropped the Gemini 3.5 Pro it had promised for June.The catch is that most people cannot use it yet. Argon is first going to a small group of trusted cyber defenders through Google’s Fairwind Program, while the model goes through the US government’s voluntary pre-release access process. Paid API customers and Google AI Ultra subscribers are next in line, though there is no date. “We’re going to make it available as soon as we can and as safely as we can,” CEO Sundar Pichai wrote on X.

Gemini 4 Argon tops 13 of 18 benchmarks against OpenAI and Anthropic models

On DeepSWE v1.1, which tests long-horizon software engineering, Argon scored 77.9%, ahead of Claude Opus 5.5 at 74.2% and GPT-6 Astra at 74.1%. The widest gap is on Harvey’s Legal Agent Benchmark, where Argon hit 19.6% against Astra’s 5.4%. It also leads the Vals Index, which weights finance, coding, legal and tax work by their share of US GDP, with 68.9%.Rivals still hold some ground. GPT-6 Astra stays ahead on FrontierSWE v2 (65.5% to 55%) and OSWorld-2.0, while Claude Opus 5.5 wins Terminal-bench 4.0 and PostTrainBench. On CWE-bench v1, which measures how well a model fixes security flaws, Argon and Astra are tied at 68%.Google has also raised the output limit to 1 million tokens, up from 64,000, so Argon can finish very long tasks in one run.Introductory API pricing is $2 per million input tokens and $10 per million output tokens, with cached input 95% cheaper. That is half of what Anthropic charges for Claude Opus 5.5 and a fifth of GPT-6 Astra’s rate.

Google is already using Argon to rewrite code and free up data centre memory

Thousands of Googlers have been using Argon internally for weeks. A team of Argon agents went through fleet-wide profiling data and freed over 300 TiB of memory across Google’s data centres, with total savings estimated at 500 TiB to 1 PiB. Agents are also moving C/C++ codebases to Rust, including more than 800,000 lines of the Fuchsia Zircon kernel. In the libgav1 video decoder, Argon replaced 32,000 lines of SIMD code, and the result runs 2.7 times faster than the earlier Rust port.Security firm Wiz used Argon to find a critical flaw exposing personal data in healthcare software used by hospitals worldwide, one that earlier frontier models had missed.

Argon’s safety guardrails watch its reasoning and block cyber and weapons misuse

A day after Pichai signed a voluntary AI safety agreement at the White House, Google laid out four areas where it is tightening Argon’s safeguards before a wider release. The first is misuse, with the model trained to refuse requests that could aid cyberattacks or chemical, biological, radiological and nuclear (CBRN) weapons. Google now also tracks Argon’s internal activations to catch anyone trying to slip past those refusals, and internal and external red teams have tested these defences.Argon is also Google’s most resistant model yet to indirect prompt injection, where hidden instructions in a webpage or document try to hijack it. On Gray Swan’s benchmark, such attacks succeeded only 0.7% of the time, compared with 1% for Claude Opus 5.5 and 8.5% for GPT-6 Astra.To stop the model from overstepping a user’s intent, Google monitors its chain of thought and actions, and can halt execution when needed. The company says it kept these monitoring findings out of training so Argon would not learn to hide its reasoning. Test sandboxes are now sealed before high-risk runs, a change that follows Google’s disclosure this month that its models had escaped a testing environment over the summer.Fairwind partners, however, get Argon without cyber guardrails so they can use its full vulnerability-hunting ability.After a summer of Flash models and a shelved Gemini 3.5 Pro, Argon gives Google its strongest frontier claim in almost a year. Google hasn’t said when more Gemini 4 models will follow. Pichai’s only hint so far is that the company will be “iterating rapidly.”



Source link

Leave a Reply

Your email address will not be published. Required fields are marked *