🔍 Read the full analysis: Can UK AISI And EvalEval Make AI Benchmark Results Easier To Reproduce? on ThorstenMeyerAI.com
Get the latest gadgets delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
TL;DR
The UK AI Security Institute is publishing selected AI benchmark results through EvalEval’s Evaluation Cards, covering five benchmarks across six frontier models plus two cyber evaluations. The release pairs scores with verification, configuration and context details so readers can see how results were produced.
The UK AI Security Institute (AISI) is publishing selected AI benchmark results through EvalEval’s Evaluation Cards, pairing scores with verification, evaluation context and configuration details that show how the results were produced. The release accompanies AISI’s paper How Inference Compute Shapes Frontier LLM Evaluation and covers five benchmarks across six frontier models, plus two cyber evaluations that use a different, partly overlapping model set.
The five benchmarks in the paper’s main experiment are HealthBench, FrontierMath, Humanity’s Last Exam, SWE-Bench Pro and Terminal-Bench 2.0. The results cover Claude Opus 4, Claude Opus 4.5, Claude Opus 4.6, GPT-5, GPT-5.2 and GPT-5.4. AISI has also shared results from two cyber evaluations, Cyber CTFs and The Last Ones, which use a model set that overlaps only partly with the main experiment — the same model list should not be assumed to apply to them.
The records are tied to AISI’s paper on how scores depend on inference-time compute and evaluation protocol. For Humanity’s Last Exam, the analysis tracks the cumulative share of attempted tasks solved within a given token count, using each task’s earliest observed success. In runs where models received correctness feedback from an oracle after each attempt, they went on to solve additional tasks as token use increased.
EvalEval describes the released records as including verified results, evaluation context and configuration information. The platform organizes benchmark metadata, evaluation-run data and model metadata into a common format. AISI says its public reporting is being made available where appropriate; the announcement does not claim that every AISI evaluation or every underlying transcript is included.
Why Setup Details Change Score Comparisons
Benchmark scores are often cited as if they measure the same thing across models, but different evaluation protocols can produce different results. The AISI paper’s Humanity’s Last Exam analysis illustrates why: outcomes shifted with inference compute and with whether models received correctness feedback between attempts. A score reported without those conditions can leave readers unsure what performance it actually represents.
Publishing results with their setup information gives researchers and practitioners a way to inspect individual evaluations and compare them with other reported runs. It can also help identify when superficially similar scores came from meaningfully different conditions. That matters for research, model development and policy work that treats evaluations as evidence about advanced AI capabilities. The records do not settle which benchmark or protocol is best, but they make some of the conditions behind a result easier to see.
The collaboration builds on earlier work between AISI and EvalEval that began at a joint workshop alongside NeurIPS 2025. EvalEval says feedback from the Institute helped shape Every Eval Ever (EEE), its shared schema for documenting evaluations, and the current release applies that shared infrastructure to publicly reported AISI methods and findings.
AISI has also worked on evaluation efficiency through OptStop, statistical rigor through HiBayES, and standardisation in areas such as transcript analysis and capability elicitation. EvalEval’s related project, Evaluation Cards, combines evaluation results with benchmark and model information. Together, these efforts address a practical reporting problem: results published across formats and outlets may omit details needed to interpret or reproduce a run, while repeating costly evaluations may not be feasible.
“AISI is using EvalEval’s infrastructure to openly share evaluation results, supporting more reproducible and verifiable evaluation science.”
— EvalEval Coalition
Coverage and Reproduction Limits
The announcement does not specify how many records or transcripts are available, which individual setup fields are present for every benchmark, or whether outside researchers have independently reproduced the results. Because it says publicly reported methods and findings are being made available where appropriate, the release should not be read as a complete archive of all AISI evaluation work.
The cyber evaluations use a different, partly overlapping model set, and the announcement does not enumerate that set. It also does not give a release date for each record or describe a process for resolving disagreements between results reported under different protocols. Those details would help readers judge the current coverage and compare the records consistently.
Broader Adoption of Every Eval Ever
EvalEval says it expects to continue standardising and sharing evaluations with AISI and other evaluation organisations. The next practical step is broader use of the Every Eval Ever schema: model developers can submit verified results, while evaluation developers can report benchmarks and run data using the schema. Researchers in evaluation, governance and policy can explore Evaluation Cards by benchmark or model and examine reporting practices across the collection.
Wider adoption could make cross-study comparisons easier, though its value will depend on the consistency and completeness of records contributors publish. No further release date or adoption milestone was specified.
Key Questions
Which benchmarks and models are covered in the main release?
Five benchmarks — HealthBench, FrontierMath, Humanity’s Last Exam, SWE-Bench Pro and Terminal-Bench 2.0 — across six models: Claude Opus 4, 4.5 and 4.6, and GPT-5, 5.2 and 5.4.
Do the cyber evaluations cover the same models?
No. Cyber CTFs and The Last Ones use a different set of models that overlaps only partly with the main experiment, and AISI has not enumerated that set.
Does the release include all AISI evaluations?
No. AISI says publicly reported methods and findings are being made available where appropriate; the announcement does not claim every evaluation or transcript is included.
Why does evaluation setup matter for interpreting scores?
AISI’s paper found that scores on Humanity’s Last Exam shifted with inference-time compute and with whether models received correctness feedback between attempts, so identical-looking scores can reflect different conditions.
Have independent researchers reproduced the results?
The announcement does not say whether outside researchers have independently reproduced the results, and it does not specify how many records or transcripts are available.
Primary source: Hugging Face · via ThorstenMeyerAI.com
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
