September 22, 2026
Benchmarks can be independently checked: UK AISI releases configurations for 5 tests
On September 22, UK AISI published results, context and configurations for 5 benchmarks in Evaluation Cards. The set includes HealthBench, FrontierMath, Humanity's Last Exam, SWE-Bench Pro and Terminal-Bench 2.0. The release covers Claude Opus 4, 4.5 and 4.6, as well as GPT-5, 5.2 and 5.4.

On June 9, Evaluation Cards in beta showed 101 955 results for 638 benchmarks. UK AISI has now added verified runs for five tests, along with the settings and context needed to reproduce them.
Your own results. Every Eval Ever accepts data in schema version 0.3.0 after local validation. Inspect AI, HELM and lm-evaluation-harness logs are converted with built-in commands, and the resulting PR to the Hugging Face datastore is reviewed by an EvalEval contributor.
Review by an EvalEval contributor remains the final step before a result is published in Every Eval Ever.
