Google releases Gemini 4 Argon: Leading in multiple benchmark tests, with output cap increased to 1 million token
Wallstreetcn
1h ago
Ai Focus
Google officially launches its flagship AI model Gemini 4 Argon, which ranked first or tied for first in 13 out of 18 benchmark tests. The output token limit has also been increased from 64,000 to 1 million. This model is initially available only to specific cybersecurity partners, and during the promotion period, it is priced at $2 per million inputs token and $10 per million outputs token.
Helpful
No.Help

Google releases its flagship AI model Gemini 4 Argon, which surpasses the top products of OpenAI and Anthropic in multiple benchmark tests, marking the return of this search giant to the forefront of competition after several months of lag. However, this highly anticipated model is currently only available to specific cybersecurity partners, and ordinary developers and enterprise users will have to wait.

On September 30th, Google officially launched its latest flagship model Gemini 4 Argon. This model ranked first or tied for first in 13 out of 18 benchmark tests, and is particularly adept at knowledge-based tasks and long-context reasoning.

In light of the widespread discussions in the tech community regarding artificial intelligence security issues in recent weeks, Google stated that its new model will initially be available only for use by specific companies and organizations focused on network security defense. The model will only be made more widely available after defenders have the opportunity to patch system vulnerabilities and fix the defects identified by the model.

Google also added that Gemini and Argon include new security measures to prevent misuse and unexpected behavior. These measures include monitoring functions that can capture instances where the models perform tasks beyond user expectations or exceed their designated boundaries, and they can also terminate the models' activities when such overstepping occurs.

Gemini, the Senior Director and Product Leader, Tulsee Doshi said, "We see that these protective measures are effective."

Knowledge work leads the way, while programming performance varies widely.

Gemini and Argon demonstrate clear advantages in benchmark tests in the field of knowledge work, but the performance in programming capabilities is more complex.

In the comparison data published by Google, Argon scored 68.9% against Vals Index, leading Opus which scored 67.0%; against Vals Finance Agent and v2, it scored 65.4%, leading Fable by about 5.1 percentage points; and against Harvey and Legal Agent Benchmark, it scored 19.6%, which is approximately three times the score of Fable at 5.1%.

In terms of programming, Argon ranked first with a score of 77.9% on DeepSWE and v1.1, leading by about 3 to 4 percentage points over Astra and Opus. Google has defined this as the new optimal level for this test.

However, on FrontierSWE and v2, Argon lags behind Astra by 55.0%, with a gap of 10.5 percentage points; on Terminal-bench 4.0, Argon also lags behind Opus's 66.4% by 57.4%, with a gap of nearly 9 percentage points.

It is worth noting that the aforementioned benchmark tests were not conducted under completely uniform experimental conditions. Taking DeepSWE as an example, Google used its own mini-swe agent framework to calculate the Argon score, while the data of competitors came from public rankings and the respective reports of each model provider. The differences in methodology mean that direct numerical comparisons need to be interpreted with caution.

A million token output limit is set, creating room for complex tasks.

An important breakthrough in technical specifications is the significant increase in the output token upper limit from the previous 64,000 to 1 million. Gemini Argon

At the technical level of AI, the difficulty of "getting the model to read 1 million token" and "getting the model to write 1 million token" are not on the same order of magnitude.

The input ( Prefill ) mainly involves parallel processing of text, and the computational cost is relatively controllable. However, the output ( Decoding ) is different: for each token generated by the model, it is necessary to recalculate using all the previously generated token , which means it is a form of autoregressive generation. Generating one million token requires maintaining the attention mechanism without collapse over an extremely long period of time, which is a significant test for both video memory, KV Cache optimization, and hardware bandwidth.

Google links this design directly to deep reasoning capabilities, believing that a larger generation space enables the model to complete more complex task chains in a single run.

Internal deployment cases disclosed by Google confirm this potential. In terms of data center memory optimization, the Argon agent has released over 300 TiB of memory by analyzing cluster performance telemetry data and implementing optimization solutions, with a potential savings estimated to range from 500 TiB to 1 PiB.

In the field of code migration, Argon has been involved in projects to migrate C/C++ code to Rust, which involve core libraries such as re2, as well as over 800,000 lines of code in the Fuchsia operating system's Zircon kernel.

During the migration of the video decoding library libgav1, the Argon intelligence replaced 32,000 lines of SIMD code, and the resulting decoder ran 2.7 times faster than the previous Rust ported version.

Google also reported that Argon assisted in optimizing the spatio-temporal resources of a subroutine in the field of quantum computing, with an improvement of up to 40%.

Network security is of utmost importance; access rights are strictly controlled.

Security defense is the focus of Argon's first batch of use cases.

Google stated that this model, after specialized training, is capable of autonomously detecting, verifying, and repairing software vulnerabilities. Trusted participants from Fairwind Program and Google's internal teams will be granted access to the original model without any additional network security safeguards.

A subsidiary of Google, Wiz, has been using Argon in its “Scan for Good” vulnerability scanning project. Google claims that this model has identified a serious flaw present in medical software used by numerous hospitals around the world, a risk that several advanced models prior to this were unable to detect.

In the internal security assessment, Argon scored 85.8% in the source code vulnerability detection test, which is 3.8% higher than Gemini and 71.0% higher than Flash Cyber; in the Wiz black box penetration testing assessment, it scored 70.9%, which is higher than the latter's 58.2%.

In the public CWE-bench v1 vulnerability repair tests, Argon tied with GPT-6 Astra at 68.0%, with Opus closely following at 67.0%.

Pricing strategy: Competitive during the promotion period, and will double thereafter.

Google has announced the pricing for Argon's API: during the promotion period, it charges $2 per million inputs of token and $10 per million outputs of token. Cached inputs enjoy a 95% discount, which translates to approximately $0.10 per million token.

After the promotion period ends, the prices will be increased to $4 per million inputs of token and $20 per million outputs of token. Both increases represent a doubling of the current rates. However, Google has not disclosed the specific end time of the promotion period.

Compared to its main competitors, Argon has a significant price advantage during the promotion period.

GPT-6 Astra and Claude Fable have a standard API price of $10 per million inputs and $50 per million outputs, which is five times the price during the Argon promotion period. The pricing for Claude Opus 5.5 is $4 per million inputs and $20 per million outputs, which is the same as the price after the promotion period for Argon.

However, for users who are looking for value for money, setting a fixed price is not an either/or situation.

GPT-6.1 and Argon have the same promotional pricing: an input of $2 results in an output of $10; Google's own Gemini currently has a pricing of $0.75 for an input of $0.75, resulting in an output of $3.75, which is even lower.

Google's benchmark tests do not include GPT-6.1 and Sol. Therefore, under the same promotional prices, the actual quality differences between the two models remain to be tested by the market.

Skip 3.5 Pro, and directly enter the Gemini 4th Era.

The release of Argon also signifies that Google has abandoned its previously planned Gemini 3.5 Pro.

Google announced at its I/O developer conference in May this year that it planned to launch Gemini 3.5 Pro in June, but this plan was never realized. Subsequently, a series of Flash models were used as replacements.

According to the analysis by the evaluation institution Vals AI (referenced by AI), this shift reflects that Gemini version 3.5 Pro is no longer sufficient to surpass its competitors in terms of performance. As a result, Google has decided to skip ahead to release Gemini version 4, in order to re-contest for the leading position in AI.

In July this year, Pichai acknowledged Google's lagging position during an investor conference call, but stated that "we are still at the forefront in many dimensions," and identified Gemini 4 as a key to catching up.

Last November, Google briefly topped the AI rankings with a new model, which triggered a "red alert" within OpenAI and prompted them to prepare urgently. Whether Argon can maintain this counterattack will still need to be verified by the market after its full rollout.

Tip
$0
Like
0
Save
0
Views 17
HKWDB reminds readers to view blockchain rationally, stay aware of risks, and beware of virtual token issuance and speculation. All content on this site represents market information or related viewpoints only and does not constitute any form of investment advice. If you find sensitive content, please click“Report”,and we will handle it promptly。
Submit
Comment 0
Hot
Latest
No comments yet. Be the first!
Related
Microsoft WinUI Completes Win11 Native Table and Chart Controls, Fixes Increased Memory Usage Bug
Microsoft has updated WinUI and Windows App SDK by adding TableView and Chart controls, enhancing the native table and chart capabilities, and fixing memory usage issues that occurred after repeatedly creating visual storyboards or continuously parsing static and theme resources. Chris Anderson, a member of Microsoft's Windows UI team, stated that performance, basic functionality, quality, and bug fixes are the priorities for WinUI.
The Block
·2026-10-02 10:54:03
0
It is reported that within less than 14 days since the launch of Apple's iPhone 18 Pro series, domestic sales have exceeded 2 million units.
According to observations by digital blogger @RD cited by IT, in less than 14 days since the launch of the iPhone 18 Pro series, sales (Sell out) have exceeded 2 million. The blogger also mentioned earlier that within 7 days of the series' launch, sales approached 1.3 million, which is approximately 115% of the iPhone 17 Pro series and 90% of the iPhone 17 series.
The Block
·2026-10-02 10:54:01
0
Signal65 reports that Qualcomm Snapdragon X2 Elite outperforms Intel Core Ultra X7 358H in aspects such as CPU and AI.
On October 1st, institution Signal65 released a report stating that Qualcomm's 18-core Snapdragon X2 Elite chip outperforms Intel's Core Ultra X7 358H in CPU, AI, productivity, and most battery life tests. In the tests, Snapdragon had an advantage in Geekbench multi-core, single-core, and several other AI tests, but Intel took the lead in office battery life tests.
The Block
·2026-10-02 10:54:00
0
MontoHealth is building the action layer of medical AI.
MontoHealth indicates that the company is researching how AI agents can surpass traditional methods and securely complete routine medical tasks. The company states that its goal is not to add another dashboard, but to act as a coordinator between existing systems, focusing on common patient interactions such as appointments, prescription refills, referrals, billing, and insurance issues.
PR Newswire
·2026-10-02 10:31:58
8
Nike's quarterly report falls short of expectations, and the company announces a new round of layoffs; annual revenue forecast significantly lower than anticipated
Nike releases financial results for the first quarter of fiscal year 2027, with revenue and net profit both falling short of expectations. The company also predicts that annual revenue will decline by more than a single digit percentage this year, far below market expectations. Additionally, Nike announced the launch of a new cost-cutting initiative named “Pace,” with the goal of saving a total of $2.5 billion by the end of fiscal year 2031.
Wallstreetcn
·2026-10-02 10:31:57
10
View More