September 24, 2026
Claude Code tops coding-agent benchmark: Opus 5.5 scores 66 points
On September 23, the Claude Code and Opus 5.5 combination scored 66 points in the Artificial Analysis index: higher than any other measured coding agent. Opus 5 scored 60 points, Claude Fable 5.1 — 62. Anthropic cut Opus pricing to $4/$20 per million input and output tokens, but one measured run costs $13.04.

Artificial Analysis
@artificialanlys
Claude Opus 5.5 became the new No. 1 in Artificial Analysis's Coding Agent Index. It improved across all three evaluations, but task cost increased. In Claude Code at maximum effort, Opus 5.5 scores 66 points in the Coding Agent Index. This is the best result we have measured. It is 6 points above Opus 5 with 60 points and 4 points above Claude Fable 5.1 with 62 points. Anthropic cut Opus prices from $5/$25 to $4/$20 per million input and output tokens. Cache read fell from $0.50 to $0.20. Even after that, task cost for Opus 5.5 is $13.04 versus $10.79 for Opus 5, because the model uses notably more tokens. Key points: ➤ Results improved across all three Coding Agent Index evaluations: Terminal-Bench 4.0 rose from 54.5% to 63.1%, DeepSWE v1.1 from 62.5% to 68.4%, and SWE-Atlas-QnA from 62.1% to 66.4%. Terminal-Bench saw the largest gain: 8.6 percentage points. ➤ The best result costs more than any other: Opus 5.5 task cost is $13.04, 21% above Opus 5's $10.79. A task uses about 15.6 million tokens versus 11.4 million for Opus 5, including 2.4 times more output tokens. ➤ Opus 5.5 extends the Pareto frontier of the Coding Agent Index by task cost. No cheaper model in our comparison matches Opus 5.5's result. The frontier has moved higher at the expensive end. Other model details: ➤ Price: $4/$20 per million input and output tokens, 20% below Opus 5. Cache read costs $0.20 per million tokens, 60% below the previous $0.50. ➤ Evaluation setup: Claude Code at maximum effort was measured on DeepSWE v1.1, Terminal-Bench 4.0, and SWE-Atlas-QnA. The Coding Agent Index gives all three evaluations equal weight.

· 2.7K views
Previously, Claude Code with Opus 5 scored 60 points and spent $10.79 per task. Opus 5.5 now scored 66, but costs $13.04: the model uses 15.6 million tokens instead of 11.4 million.
Token price fell. Anthropic cut prices from $5/$25 to $4/$20 per million input and output tokens, and cache read from $0.50 to $0.20. In practice, the math is simple: six index points cost $2.25 more for every measured task.
In the API, the model is enabled by replacing `claude-opus-5` with `claude-opus-5-5`. Thinking can no longer be disabled; an effort level must be selected. The index equally combines DeepSWE v1.1, Terminal-Bench 4.0, and SWE-Atlas-QnA: 303 tasks, with three attempts for each.
In this comparison, no cheaper model matches Opus 5.5's 66 points.
