AI Evaluation Consensus Statement

AI evaluations provide critical information for developers, deployers, and end users of AI systems. The ecosystem needs shared knowledge and consensus about what evaluations should look like, and clear guidance to make evaluations comparable and understandable.

The Consensus Statement

Based on the consensus of stakeholders and experts, we recommend evaluation practitioners consider and adopt at least these 27 practices where relevant for constructing, documenting, maintaining, using, and reporting the results of AI evaluations.

27
Consensus practices
17
Endorsing organizations
100+
Individual endorsers

A cross-sector Delphi panel of researchers, auditors, industry, standards bodies, and government identified shared, practical guidance for high-quality AI evaluations — concise community guidance in the spirit of CONSORT, PRISMA, or TRIPOD.

📋

Not a new framework

We are not creating a new evaluation framework, benchmark, or metric — only making existing cross-sector agreement visible.

⚖️

Not a mandate

Not a governance mandate, compliance regime, or audit protocol. This is guidance, not requirements, with room for context-specific judgment.

🤝

Practical consensus

Not an attempt to resolve deep methodological disputes — the focus is areas of practical agreement that make evaluations more comparable and transparent.

Consensus-level practices

These 27 practices reached our consensus threshold as strongly recommended or required across stakeholder groups.

Strong Consensus at least 2 stakeholder groups with >70% positive agreement
Broad Consensus all stakeholder groups with ≥50% positive agreement
  1. DesignStrong ConsensusBroad Consensus
    Stakeholders and roles are explained clearly
    Subitems
    • State who built and who ran the evaluation
    • State any other identified affected stakeholders for the outputs
    • Explain stakeholder/expert consultation that informed design, if any
    • Explain any normative assumptions about how/why the eval is used
  2. DesignStrong ConsensusBroad Consensus
    The report (or preregistration) explains purpose, use context, decision linkage, audience, and relation to existing evaluations
    Subitems
    • Explain the purpose and, if relevant, success objective of the evaluation
    • Explain intended use context and lifecycle placement
    • Explicitly mention any decision linkage and actions planned
    • Explain audience and interpretive boundaries
    • Report the relationship to existing evaluations and benchmark choice rationale
  3. DesignStrong ConsensusBroad Consensus
    The report provides a construct definition and scope
    Subitems
    • Explain measurement of the construct/capability and the target claim
    • Explain sub-areas/aspects of the construct, and accompanying reporting plan
    • Provide an explanation or map from construct to real-world use/applications
  4. DesignStrong ConsensusBroad Consensus
    Explain the problem framing, threat, and/or consequence model
    Subitems
    • Specify capabilities, failure modes, or threat model
    • Discuss different performance levels and outcomes
  5. DesignStrong ConsensusBroad Consensus
    Note any limits to validity and design trade-offs
    Subitems
    • Explain external validity gaps and expected limits
    • Explain representativeness and coverage, construct contamination, and item-format impacts
    • Note any internal/external validity trade-offs
  6. DesignStrong ConsensusBroad Consensus
    Specify and explain metrics, aggregation, and interpretation
    Subitems
    • Explain the metric(s) used by the evaluation
    • Explain the score meaning, and the range, including floors or ceilings, and explain any baseline(s) or natural comparison
    • Explain aggregate and subgroup metrics, and asymmetry handling if relevant
  7. DesignStrong ConsensusBroad Consensus
    Quantify uncertainty, robustness, and interpretive limits
    Subitems
    • Explain any sources of uncertainty/variation and trade-offs
    • Perform and report input variation and robustness (prompting/stochasticity)
    • Design to allow uncertainty quantification
    • Explain saturation, ceiling effects, and interpretability of above-human performance
  8. DesignStrong Consensus
    Perform and report statistical analyses
    Subitems
    • Perform power analysis or other sample size justification, and report detectable effect sizes or equivalent
    • Address multiplicity control, preregistration, and researcher degrees of freedom
    • Explain distributional assumptions, dependence structure, and model misspecification risks
  9. DesignStrong ConsensusBroad Consensus
    Explain task and item construction, and perform design validation and report
    Subitems
    • Explain the task/item development workflow, and the selection logic
    • Explain any relationship to prior evaluations/literature, and provide a rationale for any adaptation
    • Perform validation of task/item criteria, and mention any domain-expert consultation
  10. DesignBroad Consensus
    Note item sourcing, provenance, representativeness, and modality assumptions
    Subitems
    • Explain legal and ethical sourcing and permissions
    • Explain item provenance, representativeness, and coverage of items
    • Explain formats, languages, and modalities (including multimodality) and validity and robustness implications
  11. DesignStrong ConsensusBroad Consensus
    Report human-subjects ethics, oversight, and safeguards where relevant
    Subitems
    • Report any ethics/IRB approvals and oversight for human-involved components
    • Report participant safeguards and protocol transparency (when humans are involved)
  12. Before ExecutionStrong ConsensusBroad Consensus
    Provide a correctness definition and explain ground-truth validity
    Subitems
    • Provide correctness definition and ground-truth specification
    • Report validation approach and treatment of malformed/refusal outputs
    • Check and report human reference performance, and explain comparability (if relevant)
  13. Before ExecutionStrong ConsensusBroad Consensus
    Create and provide rubrics and judge-system specification
    Subitems
    • Create and provide rubric/scoring system, including construction and validation details
    • Report LLM judge(s) model/prompt details and validation
    • Report human judge(s) recruitment/demographics/compensation/consent details
  14. Before ExecutionStrong ConsensusBroad Consensus
    The evaluation performs and explains judge training and quality control processes, then reports them
  15. Before ExecutionStrong Consensus
    Perform contamination, gameability, and evaluation-awareness controls
    Subitems
    • Perform data contamination checks and exposure assessment
    • Gameability/sandbagging prevention and detection mechanisms
    • Perform evaluation-awareness checks (if relevant)
    • Report access/secrecy controls for sensitive evaluation components
  16. ExecutionStrong Consensus
    As-run logging and reproducibility provenance
  17. ExecutionBroad Consensus
    Document and report execution-time adaptation of the evaluation and elicitation controls
  18. ExecutionStrong ConsensusBroad Consensus
    Document failure patterns, run differences, and execution constraints
  19. LifecycleStrong ConsensusBroad Consensus
    Provide evaluation outputs or document artifact availability and access pathways
    Subitems
    • Licensing terms, open/modifiable availability, and model access requirements (black-box, weight access, or internal activations)
    • Provide results and replication data publicly, or explain any access alternatives if non-public
    • Provide reusable code or access conditions for third-party evaluations and future reuse
  20. LifecycleBroad Consensus
    Document governance, versioning, and re-evaluation criteria (if any exist) for future use
    Subitems
    • Provide plans for updates, deprecation, and retirement, if relevant
    • Provide version/build status and change-tracking policy
    • Explain any criteria for valid future use on new models/systems
    • Monitoring triggers and re-evaluation protocol
  21. LifecycleStrong ConsensusBroad Consensus
    Provide operational documentation, portability, and maintenance support
    Subitems
    • Include operational documents/APIs/requirements along with publication, and document any portability support
    • Provide guidelines for reuse, adaptation, and safe integration
    • The evaluation identifies maintenance ownership and feedback/contact channels, if any exist
  22. Reporting & PublicationBroad Consensus
    Include resource accounting and operational trade-offs
    Subitems
    • Report compute/sample/resource requirements for running the evaluation
    • Include cost accounting methodology and reported cost metrics (tokens, dollars, time)
    • The report describes operational trade-offs across run modes and deployment contexts
    • Explain any difference between planned resource profile and realized execution or constraints encountered
  23. Reporting & PublicationStrong ConsensusBroad Consensus
    Provide results interpretation and evidence for claims
    Subitems
    • Note any baselines and provide guidance for interpretation of results
    • Perform and report sensitivity analyses, note differences vs prior reports, if any
    • The report ties conclusions to results, relates them to decision criteria, and notes any limitations for comparisons
    • Note any disagreements or disputes over interpretation
  24. Reporting & PublicationStrong Consensus
    Report multi-model comparison standards and trend claims, if relevant
    Subitems
    • The report standardizes comparisons across reports, models/versions, and different sources of variation or uncertainty
    • Include trend analysis and future-performance projections, if possible
    • Note model selection and setup uncertainty in comparative conclusions
  25. Reporting & PublicationBroad Consensus
    Explain publication status and reporting artifacts
    Subitems
    • The report notes submission/publication status
    • Complete an evaluation card/checklist and machine-readable and human-readable reporting template
    • Include content warnings and communication safeguards (if relevant)
    • Provide a post-publication correction/update pathway
  26. Reporting & PublicationBroad Consensus
    Mention evaluation usage rules or conditions, release context, and access constraints
    Subitems
    • Report any evaluation of Pre-release, pre-deployment, or checkpoint model versions and impacts on reporting validity
    • Provide a timeline of runs/results release and recipient audiences
    • Access/publication restrictions and 'file-drawer' risks, where evaluations are not reported
    • The report notes any GUID/canary usage, or other verifiable cryptographic commitment mechanisms (e.g. hashes of non-public preregistration)
    • Include a disclosure policy or explanation for sensitive information hazards
  27. Reporting & PublicationStrong ConsensusBroad Consensus
    Provide a replication package and reproducibility guide, or explain reproducibility constraints
    Subitems
    • Make replication code and execution package available, or otherwise explain how reproducibility can be achieved
    • Ensure users can reproduce reported outputs or understand and audit irreproducible components
    • Note data retention plans and auditability of reproducibility artifacts
Practices lacking consensus 1 trending toward inclusion 1 no consensus 3 contested

Trending toward inclusion

1 practice showed positive directional support but did not yet reach the consensus threshold.

Directional (Positive) majority lean toward including this practice, but below consensus threshold
  1. DesignDirectional (Positive)
    The evaluation developers preregistered the development and prospective design plans before building the evaluation
    Subitems
    • The evaluation includes a public/verifiable design preregistration preceding the evaluation development
    • Provide a prospective plan for what the evaluation will measure, the task types and methods, the access level (black-box, weight access, or internal activations) and the evaluation creation or selection process
    • Prospectively declared task types/methods and selection process

Below threshold

4 practices did not reach consensus in either direction.

Directional (Negative) majority lean against requiring this practice
No Consensus no qualifying directional signal across stakeholder groups
  1. Before ExecutionNo Consensus
    Provide a task summary and reporting commitments before evaluation usage
    Subitems
    • Provide a summary of task inputs/types/features
    • Create and preregister a reporting plan and comparison schema
    • Explain uncertainty and disaggregation disclosure plans in the advance reporting
  2. Before ExecutionContested
    Perform prospective analysis and report plan and externally verifiable pre-run readiness checks after evaluation development, before usage
    Subitems
    • Before running the evaluation, but after development, publicly register or commit a protocol with hypotheses, methods, and decision criteria, and the relationship to any design preregistration which occurred
    • Perform pre-run verification of items, code, and metrics
    • Register and report expected performance vs. baseline (pre-run)
  3. Before ExecutionContested
    Explain data partitioning, holdouts, and revision controls
    Subitems
    • Detail pilot/tuning/evaluation/holdout separation and rationale
    • Explain any splits and construction across relevant dimensions and/or different sub-constructs for reporting
    • Revision policy after split creation and pilot feedback (if relevant)
    • Explain holdout governance for later and ongoing use
  4. Before ExecutionContested
    Pilot the evaluation and calibrate baselines
    Subitems
    • Pilot runs on pre-selected systems/configurations
    • Explain human/random/prior-model baselines for comparability
    • Pilot diagnostics for failure modes and scorer reliability
    • Pilot saturation/ceiling checks

How the consensus was developed

A preregistered, multi-round expert survey designed to surface where agreement already exists on concrete practices — with transparent reporting and no post-hoc renegotiation.

  1. Completed

    Systematic review & item development

    Broad review of evaluation frameworks, reporting checklists, and measurement standards. Candidate practices extracted, coded, and refined with expert feedback. Browse the item explorer.

  2. Completed

    Multi-sector participation

    Stakeholders across AI labs, auditors, academia, industry, standards bodies, and government contributed ratings and feedback.

  3. Completed

    Consensus Delphi rounds

    Participants rated proposed practices in areas they felt qualified to assess, with structured feedback between rounds.

  4. Completed

    Synthesis & reporting

    Results synthesized under preregistered constraints, including subgroup differences — the consensus practices above.

  5. Now

    Review & endorsement

    Stakeholders are invited to review and endorse the resulting guidance, strengthening shared expectations across the ecosystem.

Read the preregistered protocol  ·  Full process & study design

Transparency & replication

The detailed data required for reanalysis of the results is publicly available: download the replication data. No replication code is currently available; we hope to address this before finalizing the written paper.

The promised deviations log records every change between the registered protocol and the implemented study.

Endorsing organizations

Organizations that have endorsed the consensus statement.

Individual endorsers

Researchers, practitioners, auditors, policymakers, and other experts who have endorsed the consensus statement.

  • Adam Gleave Verified
    CEO, FAR.AI
  • Gary Marcus Verified
    Professor Emeritus, NYU
  • Daniel Stuart Schiff Verified
    Assistant Professor of Technology Policy, Purdue GRAIL
  • Stephen Casper Verified
    Assistant Professor, Harvard Berkman Klein Center
  • James Fox Verified
    Program Lead, Schmidt Sciences
  • Dylan Hadfield-Menell Verified
    Associate Professor, MIT
  • Colin Shea-Blymyer Verified
    Research Fellow, Center for Security and Emerging Technology
  • Leshem Choshen Verified
    Academic PI, MIT, IBM
  • Kevin Wei Verified
    Research Scholar, GovAI
  • Dr Alec Christie Verified
    Imperial College Research Fellow, Imperial College London
  • Aran Nayebi Verified
    Assistant Professor, Machine Learning Department, Carnegi...
  • Alexander Zacherl Verified
    Designer, Independent
  • Michelle Lin
    Student, Mila - Quebec AI Institute
  • Fernando Martínez-Plumed Verified
    Associate Professor, Universitat Politècnica de València
  • Alejandro Tlaie Boria Verified
    AI Policy Advisor, Pour Demain
  • Joseph Marvin Imperial Verified
    Research Fellow, Pivotal Research
  • Gathoni Ireri Verified
    Head of AI Safety Evaluations, ILINA Program
  • Vilém Zouhar Verified
    PhD, ETH Zurich
  • Ringo Lee Verified
    Commonwealth Bank of Australia
  • Vishal Goyal Verified
    Assisstant Professor, Jagannath Community College (JCC)
  • Samar Binkheder Verified
    Associate Professor, King Saud University
  • Twm Stone Verified
    Research Fellow, MATS Research
  • Cansu Yamanlar Verified
    AI Policy Researcher, CAIDP
  • Ugur Ozer Verified
    AI Governance Leader, RBC
  • Anna Sokol Verified
    PhD Student, Universoty of Notre Dame
  • Jeffrey Feng Verified
    PhD Student, UCLA
  • Swati Singh Verified
    Senior Technology Risk Professional, Financial Services
  • Burak Pişkin Verified
    AI Governance / Cybersecurity Consultant - AI Auditor, No...
  • Abbas Al Mahdi Verified
    AI Risk & Governance - Researcher & Advisor, Independent
  • Bastian Hornung Verified
    Methodology assessor, CBG-MEB
  • Sofiia Lobanova Verified
    AI Safety Researcher, Independent
  • Soung Low Verified
    Senior Model Risk Data Scientist, NatWest Group
  • Arie Horowitz Verified
    Adjunct faculty, University of Texas at Austin
  • Peter Slattery Verified
    Research Scientist, MIT AI Risk Initiative
  • Isabel Barberá Verified
    AI Product Safety Expertise Unit Lead, Dutch Coordinating...
  • Joanne Soo Min Kim Verified
    Chief Operating Officer, Trajectory Labs
  • Pratik Sachdeva Verified
    Research Scientist, UC Berkeley
  • Leon Staufer Verified
    Research Fellow, MATS
  • Dr. Erin Curfman Verified
    Instructional Design Coordinator, Lecturer, Researcher, P...
  • Noga Aharony Verified
    PhD Candidate, Columbia University
  • Kai Rawal Verified
    Student, University of Oxford
  • Anushree Chaudhuri Verified
    PhD Candidate, University of Cambridge
  • Nawaf Alampara Verified
    PhD Student, Friedrich-Schiller-Universität Jena
  • Lily Hui-Ching Wang Verified
    Professor, National Tsing Hua University
  • Juan J. Vazquez Verified
    Senior Researcher, Arb Research
  • Nicholas Diakopoulos Verified
    Professor, Northwestern University
  • Ahmed Ali Verified
    PhD Researcher, University of Zurich
  • Alexei A. Birkun Verified
    Full Professor, Medical Institute named after S.I. Georgi...
  • Paolo Giudici Verified
    Professor, University of Pavia
  • Malcolm Murray
    Research Lead, SaferAI
  • Jacy Reese Anthis
    Researcher, Sentience Institute
  • Aditya Thomas
    Research Manager, Cambridge AI Safety Hub
  • Yixiong Hao
    Co-founder, Research Manager, Second Look Research
  • Pratyush Chatterjee Verified
    Student, IIT Kharagpur
  • Godwin Abuh Faruna Verified
    Independent researcher, Independent researcher
  • Ruchira Dhar Verified
    PhD Researcher, University of Copenhagen
  • Ramesh Poluru Verified
    Senior Program Officer, The INCLEN Trust International
  • Harmonia Haomiaomiao Wang Verified
    Research Assistant, Dublin City University
  • Ihor Samokhodskyi Verified
    Founder, Policy Genome
  • Vibhu Ganesan Verified
    Product Manager, Intel
  • Dr David Gringras Verified
    Frank Knox Fellow, Harvard
  • Anna M. Pastwa Verified
    Director of AI Deployment Assurance, Standard Chartered Bank
  • Lauren Damme Verified
    Senior Fellow, Data Foundation
  • Seibu Mary Jacob Verified
    Senior Lecturer in Engineering Mathematics/Analytical Met...
  • Tuesdasy Verified
    Director of Research, ARTIFEX Labs / MLCommons AIRR Group
  • Naveen Goud Bobburi Verified
    Chief Manager, ICICI Bank
  • Himanshu Joshi Verified
    Lead Researcher, SafeAlign AI
  • Heera Sharma Verified
    NA, NA
  • David Manheim Verified
    Technion - Israel Institute of Technology
  • Paulo Bova
    Modeling Cooperation, Teesside University
  • Sriram Krishnan
    Founder, Finevals
  • Jose H. Orallo Verified
    Director of Research, Leverhulme Centre for the Future of...
  • Yixiong Hao Verified
    Co-director, Georgia Tech AI Safety Initiative
  • Irving Torres
    AI Evals Consensus
  • David Williams-King Verified
    Research Manager, ERA
  • Himanshu Joshi Verified
    CEO, Collective Human + Machine Intelligence (COHUMAIN) Labs
  • Deep Joshi
    Student, Nirma University
  • Naveen Raman
    PhD Student, Carnegie Mellon University
  • Louis Yiven Zhu
    Researcher, Oxford Internet Institute
  • Mahalakshmi Kalyanasundaram
    Forward Deployment Engineer, cohumain.ai
  • Atrisha Sarkar
    Assistant Professor, Western University, Canada.
  • Shivani Shukla
    Chieft Scientist, COHUMAIN Labs
  • Elliot Bearden
    Research Engineer, Independent
  • Henry Papadatos
    Executive Director, SaferAI
  • Sree Harsha Nelaturu
    PhD Student, Zuse Institute Berlin
  • Kevin O'Shaughnessy
    Research Engineer, Independent, University of Surrey
  • KSHITIJ JAYESH THAKKAR
    Founder, TraceVerse
  • Vojtěch Kovařík
    Postdoctoral Researcher, Czech Technical University, Char...
  • Rustam Shariq Mujtaba
    East Asia and Southeast Asia Regional Youth Champion, Par...
  • Ramakrishnan Veeramony
    Board AI Advisor, MIT AI Risk Initiative
  • Hicham Yezza
    Evaluation Lead, Responsible AI, BBC
  • Amanda Starling Gould, PhD
    Associate in Research, Duke University
  • Pooja Bhatia
    Product manager, HCL
  • Kenneth Enevoldsen
    Assistant Professor, Aarhus University
  • Lawrence Lwanji
    CEO, NGO
  • Andrea Loehr Verified
    AI Evals Consensus
  • Drishti Sharma
    AI Evals Consensus
  • Yilin Huang
    AI Evals Consensus
  • Nidhi Sakpal
    Algoverse
  • Ishan Khire
    AI Evals Consensus

Expert panel

Delphi panel participants represented a cross-sector group of organizations.

  • AI Security Institute
  • AVERI
  • Carnegie Mellon University
  • Center for Open Science
  • Center for Security and Emerging Technology
  • Centre for Governance of AI
  • ETH Zurich
  • European AI Office
  • Google DeepMind
  • Harvard University
  • Hugging Face
  • IBM Research
  • Imperial College London
  • Intel
  • METR
  • MIT CSAIL
  • NYU
  • OpenAI
  • Oxford Internet Institute
  • RAND Center on AI, Security, and Technology
  • Schmidt Sciences
  • Tsinghua University
  • UC Berkeley

Partial list of affiliations represented during the Delphi process. Participation does not imply endorsement of the final results.