September 19, 2026
Models can be measured on your own repository: Vals Smith offers 120 free credits
Vals Smith builds a benchmark from your GitHub repository: merged pull requests become tasks with hidden tests.

Your own repository can now become a benchmark: the service takes your merged pull requests and hides the tests from the model.
Vals AI launched this setup. The company raised $40 million from a16z on 13.08.2026 at a $400 million valuation, and its leaderboards are openly available: 30+ benchmarks and more than 50 assessed models, with no registration.
How to build your own measurement. Sign in with GitHub, then choose a public or private repository, code models or agents, and a harness. The service returns a pass rate for each task. At launch, it provides 120 free credits (Vals AI blog, 13.08.2026).
**One measurement does not answer your question.** For maintaining someone else’s code, Claude Opus 5 has the best result — 28,53% (Vals data, 17.09.2026). The test has 100 applications and up to 10 related changes in a row on top of their own code. For greenfield builds, the same models score 90% (Vibe Code Bench v1.1, 15.09.2026). A number from someone else’s leaderboard does not answer the question about your code.
Vals does not publish its test materials so models cannot train on future questions. Vals Index 2.0 from 15.09.2026 combines 7 benchmarks across finance, coding, and legal work, weighting sectors by their share of US GDP.
The next Vals Index update will show whether anyone can catch up with Claude Fable 5.1.
