OFICIAL Hugging Face Blog

How UK AISI and EvalEval Are Making Benchmark Results Reproducible

What happened
Based on Hugging Face Blog · Sep 22, 2026

The UK AI Security Institute (AISI) and EvalEval are collaborating to publish reproducible AI evaluation results using EvalEval’s shared infrastructure, aiming to improve transparency and comparability in AI benchmarking.

How UK AISI and EvalEval Are Making Benchmark Results Reproducible
Hugging Face Blog — Hugging Face
Key points
·
AISI and EvalEval collaborate to publish reproducible AI evaluation results using shared infrastructure and the Every Eval Ever schema
·
AISI’s Evaluation Cards release includes verified results for six frontier models and two cyber evaluations with full configuration details
·
The partnership emphasizes transcript-level transparency to improve analysis and diagnosis of evaluation results across benchmarks
Key numbers
·
6, GPT-5, GPT-5.
·
2, and GPT-5.

The EvalEval Coalition and the UK AI Security Institute (AISI) have launched a collaboration to openly share AI evaluation results, enhancing reproducibility and verifiability in evaluation science. This initiative builds on prior joint research, including work initiated at a NeurIPS 2025 workshop, which contributed to the development of the Every Eval Ever (EEE) schema. The partnership now applies this shared infrastructure to practical deployment, addressing challenges in inconsistent reporting formats and limited reproducibility across AI evaluations.

AISI is using EvalEval’s Evaluation Cards platform to publish verified evaluation results, context, and configuration details for five benchmarks in its paper, *How Inference Compute Shapes Frontier LLM Evaluation*. The release covers six frontier models—Claude Opus 4, Claude Opus 4.5, Claude Opus 4.6, GPT-5, GPT-5.2, and GPT-5.4—alongside results from two cyber evaluations, Cyber CTFs and The Last Ones. This transparency allows researchers to examine individual studies and compare findings across the ecosystem more reliably.

The collaboration emphasizes transcript-level transparency to support not only reproducibility but also deeper analysis and diagnosis of evaluation results. By providing verified reference points, AISI’s releases help researchers understand how evaluation setup choices, such as inference-time compute and protocol variations, influence reported performance. This approach addresses gaps in current evaluation reporting practices, where lack of detail often obscures meaningful comparisons between studies.

The EvalEval Coalition, a research community, aims to standardize evaluation science through projects like Every Eval Ever and Evaluation Cards, which consolidate benchmark metadata, evaluation-run data, and model metadata into interpretable records. AISI, as part of the UK government’s Department for Science, Innovation and Technology, focuses on equipping governments with scientific insights into AI risks, conducting research, and developing infrastructure to inform policy and mitigation strategies.

Original source → Deals on Clipraptor.com →