September 28, 2026
Evaluating AI agents for code security is easier: Cyber Index puts leaders at 56 points
On Sep 28, Artificial Analysis put Grok 4.7 and MiMo-V2.6-Pro at the top of the Cyber Index with 56 points each. GPT-6 Luna scored 53, GLM-5.3-Flash 50, and Muse Spark 1.3 44. The index combines three open benchmarks where AI agents find and patch vulnerabilities.

Artificial Analysis
@artificialanlys
Introducing the Artificial Analysis Cyber Index and the Artificial Analysis Cyber Index Alliance: a new standard for evaluating AI models in enterprise security. The Artificial Analysis Cyber Index Alliance brings together industry partners to create a new standard for evaluating AI models on enterprise security tasks. The alliance launches alongside the Artificial Analysis Cyber Index, which combines three open benchmarks from partners and tests how AI agents find and fix vulnerabilities. Models are getting better at cyberattack tasks, making it more important for AI labs and companies to understand how they perform in cyber defense and which models deliver the strongest results. Today, we are launching the Cyber Index Alliance with Collinear AI, IBM, NVIDIA, and Vercel. Artificial Analysis Cyber Index benchmarks: ➤ CWE-Bench-AA from Collinear AI tests auditing and patching: 120 closed tasks across all ten categories of the 2025 OWASP Top 10 for C/C++, Go, Java, JavaScript/TypeScript, Python, and Rust. ➤ DeepsecBench-AA from Vercel isolates vulnerability discovery: an agent receives a codebase and a budget, must find every vulnerability, and is compared against a reference set of findings from security specialists. Real findings earn points, while safe code incorrectly marked as vulnerable loses points. ➤ CyberGym-E2E-AA tests the full path: find a memory bug, write a proof-of-concept that triggers a crash, then fix it so the crash can no longer be reproduced. Key results: ➤ Grok 4.7 (xhigh) and MiMo-V2.6-Pro lead the Cyber Index with 56 points. They are followed by GPT-6 Luna (max) with 53, GLM-5.3-Flash with 50, and Muse Spark 1.3 (xhigh) with 44. ➤ Safety-based refusals hold back several top-tier models: GPT-6 Sol (max), GPT-6 Astra (max), Claude Opus 5.5 (max with a fallback model), Claude Fable 5.1 (max with a fallback model), and Gemini 3.8 Flash (high) refuse tasks accounting for 32–38% of the Cyber Index. Despite their agentic coding capabilities, they trail the leaders by 19–31 points. Most of the gap comes from CyberGym-E2E-AA: GPT-6 Sol and GPT-6 Astra refused every task, Claude Opus 5.5 refused 98%, and Claude Fable 5.1 refused 99%.

· 34.9K views
CWE-Bench-AA tests auditing and patching across 120 closed tasks in ten OWASP Top 10 categories. DeepsecBench-AA assesses whether an agent finds every vulnerability in a codebase without flagging safe code as dangerous. CyberGym-E2E-AA requires agents to find a memory bug, reproduce the crash, and write a patch for 131 C/C++ projects in 90 minutes.
Now there is a shared score. Cyber Index combines the results of the three tests with equal weighting. Safety-based refusals leave GPT-6 Sol, GPT-6 Astra, Claude Opus 5.5, Claude Fable 5.1, and Gemini 3.8 Flash 19–31 points behind the leaders: in CyberGym-E2E-AA, these models refused at least 98% of tasks.
Vercel already provides a command for running your own scan through Deepsec: `pnpm deepsec process --project-id my-app --agent pi --model xai/grok-4.5`.
Artificial Analysis runs all three evaluations on the open Stirrup harness, installed with `pip install stirrup`.
