September 22, 2026
Grok 4.7 beats GPT-5.6 Sol among coding agents: 56 points in the Artificial Analysis index
The Grok 4.7 and Grok Build combination scored 56 points in Artificial Analysis's Coding Agent Index, 9 more than the previous version.

Artificial Analysis
@artificialanlys
Grok 4.7 scores 46 in the Artificial Analysis Intelligence Index, putting SpaceXAI in the top four AI labs. Its Coding Agent Index results also improved: the model overtook GPT-5.6 Sol. Grok 4.7 adds +2 points over Grok 4.6 in the Intelligence Index, with strong results on agentic knowledge-work tasks. We evaluated the new model at the xhigh reasoning level. Congratulations to @SpaceXAI and @ElonMusk on the release! Key points: ➤ Grok 4.7 reaches the frontier of agentic knowledge work: +111 Elo over Grok 4.6 (high) on AA-Briefcase, our private benchmark for long-horizon agent tasks. Its 1657 Elo puts the model alongside Claude Opus 5 and Claude Fable 5.1 at the very frontier. On GDPval-AA, it scores 1695 Elo, +90 above Grok 4.6 (high). ➤ A jump for coding agents: Grok 4.7 (xhigh) with Grok Build scores 56 in the Artificial Analysis Coding Agent Index, +9 over Grok 4.6 (xhigh). Among models in their native harnesses, Grok 4.7 + Grok Build now ranks 4th, behind only Claude Fable 5.1, GPT-6 Astra, and Claude Opus 5. ➤ Elsewhere, changes are selective: outside agentic knowledge work, Grok 4.7 matches Grok 4.6 (high) on other Intelligence Index tasks. It gains on Terminal-Bench 4.0 (+4.5 pp) and GDP.pdf (+3.0 pp), while slipping on AA-LCR (−3.7 pp) and AutomationBench-AA (−1.1 pp). ➤ High token consumption: Grok 4.7's gains come at a costly token price. Grok 4.7 (xhigh) uses about 81k output tokens per Intelligence Index task, while Grok 4.6 (high) uses 36k and GPT-6 Astra (max) 27k. That is 125% and 196% more. Other model details: ➤ A 500k-token context window, as with Grok 4.6 ➤ $2/$6 per 1M input/output tokens, with cache hits discounted to as low as $0.50 per 1M, as with Grok 4.6 ➤ The reasoning level is configurable from low to xhigh. We evaluated it at xhigh.

· 21.5M views
Grok pushed GPT-5.6 Sol aside in a benchmark of coding agents. Only three remain ahead.
Artificial Analysis tested Grok 4.7 at the maximum reasoning level on 21.09.2026 and gave the combination with Grok Build 56 points in the Coding Agent Index. Only Claude Fable 5.1, GPT-6 Astra, and Claude Opus 5 ranked higher.
The price remains unchanged. SpaceXAI charges $2 per million input tokens and $6 per million output tokens, the same as Grok 4.6 cost. Rates double for prompts longer than 200 thousand tokens.
The bill will still rise. On one index task, Grok 4.7 burns through about 81 thousand output tokens; Grok 4.6 used 36 thousand. The per-token rate is the same, but the model spends more than twice as many tokens.
You can install the combination with one command: `curl -fsSL https://x.ai/cli/install.sh | bash` on macOS and Linux, then run `grok` in the project folder with the `-m grok-4.7` flag. The model appeared in Cursor on release day, where the window is 256k by default and 500k in long-context mode.
In Artificial Analysis's overall ranking, the new model scored 46 points versus 44 for Grok 4.6, putting SpaceXAI among the top four labs. The gains are uneven: up 4.5 percentage points on Terminal-Bench 4.0, down 3.7 on long-context AA-LCR.
The Grok 4.7 Fast variant responds twice as fast at twice the price, but is not offered through the public API.
