Support concise, community-endorsed guidance for how AI evaluations should be conducted and reported. Add your name as an individual or on behalf of your organization.
AI evaluations provide critical information for developers, deployers, and end users of AI systems. The ecosystem needs shared knowledge and consensus about what evaluations should look like, and clear guidance to make evaluations comparable and understandable.
The Consensus Statement
Based on the consensus of stakeholders and experts, we recommend evaluation practitioners consider and adopt at least these 27 practices where relevant for constructing, documenting, maintaining, using, and reporting the results of AI evaluations.
Provide a replication package and reproducibility guide, or explain reproducibility constraints
Subitems
Make replication code and execution package available, or otherwise explain how reproducibility can be achieved
Ensure users can reproduce reported outputs or understand and audit irreproducible components
Note data retention plans and auditability of reproducibility artifacts
Practices lacking consensus1 trending toward inclusion1 no consensus3 contested
Trending toward inclusion
1 practice showed positive directional support but did not yet reach the consensus threshold.
Directional (Positive) majority lean toward including this practice, but below consensus threshold
DesignDirectional (Positive)
The evaluation developers preregistered the development and prospective design plans before building the evaluation
Subitems
The evaluation includes a public/verifiable design preregistration preceding the evaluation development
Provide a prospective plan for what the evaluation will measure, the task types and methods, the access level (black-box, weight access, or internal activations) and the evaluation creation or selection process
Prospectively declared task types/methods and selection process
Below threshold
4 practices did not reach consensus in either direction.
Directional (Negative) majority lean against requiring this practice
No Consensus no qualifying directional signal across stakeholder groups
Before ExecutionNo Consensus
Provide a task summary and reporting commitments before evaluation usage
Subitems
Provide a summary of task inputs/types/features
Create and preregister a reporting plan and comparison schema
Explain uncertainty and disaggregation disclosure plans in the advance reporting
Before ExecutionContested
Perform prospective analysis and report plan and externally verifiable pre-run readiness checks after evaluation development, before usage
Subitems
Before running the evaluation, but after development, publicly register or commit a protocol with hypotheses, methods, and decision criteria, and the relationship to any design preregistration which occurred
Perform pre-run verification of items, code, and metrics
Register and report expected performance vs. baseline (pre-run)
Before ExecutionContested
Explain data partitioning, holdouts, and revision controls
Subitems
Detail pilot/tuning/evaluation/holdout separation and rationale
Explain any splits and construction across relevant dimensions and/or different sub-constructs for reporting
Revision policy after split creation and pilot feedback (if relevant)
Explain holdout governance for later and ongoing use
Before ExecutionContested
Pilot the evaluation and calibrate baselines
Subitems
Pilot runs on pre-selected systems/configurations
Explain human/random/prior-model baselines for comparability
Pilot diagnostics for failure modes and scorer reliability