September 23, 2026
The more bugs an agent fixes, the costlier the run: 37 Bug Hunt Bench models
As of September 23, Bug Hunt Bench has collected 133 entries from 37 models: the cost of a run rises exponentially with the number of bugs fixed.

Previously, the Score vs cost chart required comparing the number of bugs fixed and cost row by row. Now Pavel Khuryn used ChatGPT to overlay a trend line: a model above it fixes more bugs for its money.
In the July wave, GPT-5.6 Luna on max fixed 33 of 105 bugs for $1.80. Fable 5 fixed 24 bugs for $68.07.
Same route. Each model gets one prompt and one run in each of two repositories with 105 pre-seeded bugs. An independent judge blindly compares the diff against a hidden answer key.
Cost on the board may be an actual bill, a price-list calculation, a lower estimate, or zero for free. The author asks readers not to rank figures of different types as comparable dollars.
The author adds models after a GitHub Issue that includes the model’s name, its own CLI, and launch instructions.
Source
