October 1, 2026
Vulnerability hunting gets cheaper: GPT-6 Luna costs up to 100 times less than Grok 4.7
Artificial Analysis estimates as of Oct 1 put the cost of 100 bug hunts with GPT-6 Luna and MiMo-V2.6-Pro at about $20.

Artificial Analysis
@artificialanlys
It is hard to restrict offensive use without blocking defense: finding and confirming a vulnerability requires the same steps, whether the goal is to exploit it or fix it. CyberGym-E2E-AA tests models' ability to defend code, from vulnerability discovery to a fix, on tasks involving memory errors. For some of the most intelligent models, safety restrictions block responses on 85% of tasks or more. The good news: some of the most capable models are also the most cost-effective. With GPT-6 Luna or MiMo-V2.6-Pro, you can perform roughly 100 bug hunts in a codebase of over a million lines for about $20. That is up to 100 times cheaper per task than the next most capable model, Grok 4.7.

· 36.2K views
Some models refuse to look for vulnerabilities more often than they fail to find them: on CyberGym-E2E-AA, six models block at least 98% of tasks.
To fix a vulnerability, an AI agent has to reproduce the crash using the same actions an attacker would take. Artificial Analysis tests the full path from independent discovery to a patch on tasks involving memory errors.
The cost of a bug hunt. According to Artificial Analysis estimates from Oct 1, GPT-6 Luna and MiMo-V2.6-Pro can perform roughly 100 bug hunts in a codebase of over a million lines for $20. Each task costs up to 100 times less than with Grok 4.7, the next model by task success rate.
| Model | Solved on the first attempt | | --- | --- | | MiMo-V2.6-Pro | 78.6% | | GPT-6 Luna (Max) | 77.9% | | Grok 4.7 (Xhigh) | 74.0% |
Artificial Analysis results as of Oct 1, 2026. The set contains 131 tasks, one per project. The original CyberGym-E2E contains 920 vulnerabilities across 139 projects. The full list of selected tasks and CPU and RAM requirements are published in the methodology.
Independent discovery is markedly different from fixing a bug with a supplied example. In a Berkeley RDI study from Jun 18, Claude Opus 4.5 with Claude Code fixed 82.3% of tasks when given a PoC and crash log. When searching independently, its success rate was 19.2%. Each task had a budget of $10 and 90 minutes.
Artificial Analysis also allows 90 minutes per task. To count as a success, the agent must reproduce the crash, fix it with a patch and pass the project's tests. Whether the patch fixes the original vulnerability is checked separately; that check is not included in the overall success rate.
According to Artificial Analysis data from Sep 28, the following models refuse to perform at least 98% of tasks:
- GPT-6 Astra; - GPT-6 Sol; - Claude Fable 5.1; - Claude Opus 5.5; - Qwen3.8 2.4T A95B; - Qwen3.8 27B.
Among the remaining models, 42% of attempts hit the 90-minute limit without producing inputs that reproduce the crash.
How to run it. CyberGym-E2E is available on GitHub. The Berkeley RDI README includes dependency installation instructions. Download the dataset with `hf download sunblaze-ucb/cybergym-e2e --repo-type dataset --local-dir data/` and the Docker images with `python scripts/pull_images.py`.
Run a single bug hunt and fix with `python scripts/run_agent.py curl/arvo_66012 --mode e2e`. The `--mode patch-only` mode gives the agent a ready-made PoC and crash log to test its ability to fix the bug without finding it independently.
For its measurements, Artificial Analysis uses Stirrup, an open-source agent runtime. Install the base package with `pip install stirrup`, or the version with all optional components with `pip install 'stirrup[all]'`.
In 31% of successful attempts, agents fix a different real vulnerability instead of the one the task was created for.