September 19, 2026
AI-agent leaderboard now counts refusals: Fable 5.1 posted 8,8% fallback in Claude Code
On September 18, 2026, Artificial Analysis added the Safety Refusal Rate metric to Coding Agent Index v1.5. It counts attempts in which a model or provider refused to start or continue a task. Claude Fable 5.1 showed the highest fallback share: 8,8% of index weight in Claude Code and 7,1% in Devin Fusion.

Artificial Analysis
@artificialanlys
Safety refusal reporting is now available in the Artificial Analysis Coding Agent Index. In the latest Coding Agent Index v1.5, we added safety refusal reporting to explain model behavior and score differences. A safety refusal occurs when a provider or model refuses to start or continue a task for safety reasons. An agent can switch to another model to continue working or stop the attempt with a block. Claude Fable 5.1 had the highest fallback share in Claude Code and Devin Fusion: these attempts accounted for 8,8% and 7,1% of index weight, respectively. These results therefore include the performance of models used during fallback. Observed shares may be affected by refusal variance, context accumulation in the harness, effort settings, and retry strategies.

· 44.3K views
Before v1.5, the index did not separate safety refusals into a standalone metric. Safety Refusal Rate now shows the share of attempts in which a model or provider did not start or continue a task. The index aggregates results from 303 tasks across three equally weighted benchmarks: DeepSWE v1.1, Terminal-Bench 4.0, and SWE-Atlas-QnA. Each task is run three times.
How the result is calculated. On fallback, Claude Code and Devin Fusion can switch to another model or continue with the same model, and the attempt receives its normal score. For Claude Fable 5.1, these attempts accounted for 8,8% of the weight in Claude Code and 7,1% in Devin Fusion. A blocked attempt receives zero, while 2% blocked safety refusals can reduce the final index by at most 2 points.
The final v1.5 score already includes the performance of models used during fallback.
