October 1, 2026
Who finished the code after a refusal? Coding Agent Index now shows the fallback model
An agent's ranking includes the fallback model's work if it completed the task after the primary model refused.

Artificial Analysis
@artificialanlys
The Artificial Analysis Coding Agent Index now shows when safety refusals occur and which models take over from those that refuse. We've added two ways to explore the results: ➤ Refusal timing: see whether the agent refused immediately based on the task prompt or later, after it had started working. ➤ Fallback model: see which model the agent switched to after a refusal, as well as attempts that were halted. Claude Code with Sonnet 5.5 (max) currently ranks first in the Index. Its safety refusal rate is 4.5%, almost half the 8.9% rate for Claude Code with Opus 5.5 (max). Around 94% of Sonnet 5.5 refusals occurred after the first turn. After a refusal, the agent almost always switched to Opus 4.8. We see different fallback patterns depending on the model and agent configuration. With Fable 5.1, Claude Code mainly switched to Opus 4.8, while Opus 5 accounted for a much larger share of switches for Devin Fusion in the Index. Both views are available for the overall Index and each benchmark. The provider configures safety refusals and fallback switches and may change these settings over time. These rates reflect the behavior we recorded during testing.

· 3.5K views
With Claude Code running Sonnet 5.5 (max), almost 94% of refusals occurred after work had already begun. After a refusal, the agent almost always switched to Opus 4.8.
Artificial Analysis previously showed the refusal rate. Since Oct 1, you can see when an agent refused and which model took over the task. When choosing an AI coding agent, read the final score as the result of the whole setup, including the fallback model.
Even the leader switches. Claude Code with Sonnet 5.5 (max) ranks first in the Index and refuses almost half as often as the configuration with Opus 5.5 (max). Artificial Analysis recorded these results as of Oct 1:
| Model in Claude Code | Refusal rate | | --- | --- | | Sonnet 5.5 (max) | 4.5% | | Opus 5.5 (max) | 8.9% |
The fallback model is chosen according to the provider's settings. With Fable 5.1, Claude Code mainly switched to Opus 4.8. In Devin Fusion, Opus 5 accounted for a noticeably larger share of switches.
In the v1.5 results from Sep 18, Fable 5.1 had the highest rate of switches to fallback models. These switches were already included in the final scores:
| Agent with Fable 5.1 | Share of Index weight involving fallback switches | | --- | --- | | Claude Code | 8.8% | | Devin Fusion | 7.1% |
Where to see the switches. In the Coding Agent Index, open Safety Refusals → Safety Refusal Rate. The Timing toggle separates refusals based on the task prompt from refusals after work has begun. Fallback model shows halted attempts under the Blocked label and switches to Claude Opus 4.8, Claude Opus 5 or GPT-5.6 Sol.
Both views are available for the overall Index and individual benchmarks:
- DeepSWE v1.1. - Terminal-Bench 4.0. - SWE-Atlas-QnA.
Artificial Analysis includes both halted attempts and switches to fallback models in Safety Refusal Rate. Each benchmark receives equal weight, while retries that were superseded by other attempts are excluded from the calculation.
An agent receives 0 points for Blocked attempts, while tasks completed by a fallback model are scored normally. A Blocked share of 2% means losing at most 2 Index points.
Providers can change refusal and fallback model settings, so these rates apply to the recorded runs.
