September 30, 2026
Gemini 4 Argon matches GPT-6 Astra's score at a lower cost: 60% of the task price with a launch discount
Google matched OpenAI in the Artificial Analysis rankings, while an introductory discount lowered the cost of a benchmark task.

Artificial Analysis
@artificialanlys
Google's new Gemini 4 Argon has matched GPT-6 Astra in the Artificial Analysis Intelligence Index, and with discounts, a task costs 60% of what it costs with Astra. Google is back among the top three labs for intelligence performance. Gemini 4 Argon is @GoogleDeepMind's first closed model above the Flash tier in more than 7 months. At the high reasoning level, the highest available, it scores 53 points in the Artificial Analysis Intelligence Index. That matches GPT-6 Astra (max, 53) and is 1 point ahead of GPT-6.1 Sol (max, 52). The gains came from fewer hallucinations and stronger agent capabilities. With the current 50% pricing discount and an increased 95% discount on cached input, an Intelligence Index task costs $1.99 with Gemini 4 Argon. That is 60% of the cost of GPT-6 Astra (max), but 2.7 times as much as GPT-6.1 Sol (max). Once the promotion ends, the cost will rise to $3.98, roughly 1.2 times the cost of GPT-6 Astra (max). Gemini 4 Argon is currently rolling out to select users. There is no public launch yet. The 50% discount is an introductory offer. Google has not yet confirmed when it will end. Key test results for Gemini 4 Argon at the high reasoning level: ➤ Google is back among the top three labs for intelligence performance: Gemini 4 Argon (high) scores 53 points in the Artificial Analysis Intelligence Index. That matches GPT-6 Astra (max, 53) and is 1 point ahead of GPT-6.1 Sol (max, 52). The score is 23 points higher than Google's previous model outside the Flash tier, Gemini 3.1 Pro Preview (30), and 12 points higher than Gemini 3.8 Flash (high). ➤ The introductory 50% discount makes Gemini 4 Argon competitive on cost per task: with the current discount, an Intelligence Index task costs $1.99 with Gemini 4 Argon (high). That is 60% of the cost of GPT-6 Astra (max, $3.26) at a comparable intelligence level. The savings come from lower token prices, rather than lower token usage. Gemini 4 Argon uses an average of 62,000 output tokens per task, while GPT-6 Astra (max) uses 27,000. Google has not yet confirmed when the promotion will end. At standard prices, the cost per task will rise to $3.98. ➤ The model performs better on agent tasks: this has historically been a weakness for Gemini, but Gemini 4 Argon improved its results on these benchmarks. It ranks first in AutomationBench-AA at 77.5%, 6 percentage points above Claude Sonnet 5.5 (max, 71.3%). In Terminal Bench 4, Gemini 4 Argon scores 57%, a gain of 53 percentage points over Gemini 3.1 Pro Preview. Only Claude Sonnet 5.5 (max, 64%), Claude Opus 5.5 (max, 60%) and GPT-6 Astra (59%) score higher. In AA-Briefcase, the model reaches 1,494 Elo. This result was driven by a 65% evaluation criteria completion rate, the highest we have recorded. Its analysis quality and presentation scores are lower, however: 1,576 Elo and 1,308 Elo, respectively. ➤ The lowest hallucination rate among leading models: in AA-Omniscience, Gemini 4 Argon fabricates answers in 15% of cases. That is the lowest rate among models scoring at least 45 points in the Intelligence Index. By comparison, GPT-6 Astra (max) has a rate of 51%, and GPT-6.1 Sol (max) has a rate of 54%. Argon is much more likely to admit it does not know the answer instead of making an incorrect guess. Gemini 4 Argon's share of correct answers is 50%. That is 5 percentage points below Gemini 3.1 Pro Preview and 13 points below GPT-6 Astra (max, 63%). Despite this somewhat lower accuracy, Argon's overall AA-Omniscience score is 42 points, keeping it on par with GPT-6 Astra (43) and GPT-6.1 Sol (42). Main model specifications: ➤ Context window: 1 million tokens. ➤ Multimodality: accepts text, images, video and speech; responds with text. ➤ Pricing: standard pricing is $4 per 1 million input tokens and $20 per 1 million output tokens. A 50% discount is currently in effect, reducing the prices to $2 and $10. Cached input tokens receive a 95% discount: with the introductory offer, they cost $0.10 per 1 million. Gemini 3.8 Flash offered a 90% discount on cached input. ➤ Long Decode Continuation: we tested Gemini 4 Argon with a new Gemini API feature that pauses long responses and continues them in subsequent calls. This lets the model reason for up to 1 million output tokens without the request timing out.

· 11.2K views
Gemini 4 Argon scored 53 points in the Artificial Analysis Intelligence Index as of Sep 30. GPT-6 Astra earned the same score.
The gap has narrowed. The previous Gemini model above the Flash tier, Gemini 3.1 Pro Preview, scored 30 points. Argon gained 23 points and outscored Gemini 3.8 Flash by 12. Artificial Analysis attributes the gains to fewer fabricated answers and improvements on AI agent tasks.
Discounted pricing. Google cut Argon's launch prices by 50%. In Artificial Analysis testing, a task costs less than with GPT-6 Astra but more than with GPT-6.1 Sol.
| Model and mode | Cost per Intelligence Index task as of Sep 30, 2026 | | --- | --- | | Gemini 4 Argon, high, discounted | $1.99 | | Gemini 4 Argon, high, standard pricing | $3.98 | | GPT-6 Astra, max | $3.26 | | GPT-6.1 Sol, max | Discounted Argon costs 2.7 times as much |
Argon's savings come from token pricing. It uses an average of 62,000 output tokens per task, compared with Astra's 27,000. Once the introductory offer ends, Argon's cost per task will double and exceed Astra's.
Argon took first place in AutomationBench-AA with a score of 77.5%. In Terminal Bench 4, the model scored 57%, compared with 4% for Gemini 3.1 Pro Preview. In this test of terminal work, Argon trailed Astra and two Claude models.
Argon guesses less often when it does not know the answer. In AA-Omniscience, its hallucination rate was 15%, compared with Astra's 51%, but its share of correct answers was also lower: 50% versus 63%.
There is no public launch yet. Google is beginning to give the first group of cyberdefense specialists access to Argon through the Fairwind Program.
Google promises access for developers, businesses and everyday users in the next phase.
