AI evaluations provide critical information for developers, deployers, and end users of AI systems. The ecosystem needs shared knowledge and consensus about what evaluations should look like, and clear guidance to make evaluations comparable and understandable.
Based on the consensus of stakeholders and experts, we recommend evaluation practitioners consider and adopt at least these 27 practices where relevant for constructing, documenting, maintaining, using, and reporting the results of AI evaluations.
A cross-sector Delphi panel of researchers, auditors, industry, standards bodies, and government identified shared, practical guidance for high-quality AI evaluations — concise community guidance in the spirit of CONSORT, PRISMA, or TRIPOD.
We are not creating a new evaluation framework, benchmark, or metric — only making existing cross-sector agreement visible.
Not a governance mandate, compliance regime, or audit protocol. This is guidance, not requirements, with room for context-specific judgment.
Not an attempt to resolve deep methodological disputes — the focus is areas of practical agreement that make evaluations more comparable and transparent.
These 27 practices reached our consensus threshold as strongly recommended or required across stakeholder groups.
1 practice showed positive directional support but did not yet reach the consensus threshold.
4 practices did not reach consensus in either direction.
A preregistered, multi-round expert survey designed to surface where agreement already exists on concrete practices — with transparent reporting and no post-hoc renegotiation.
Broad review of evaluation frameworks, reporting checklists, and measurement standards. Candidate practices extracted, coded, and refined with expert feedback. Browse the item explorer.
Stakeholders across AI labs, auditors, academia, industry, standards bodies, and government contributed ratings and feedback.
Participants rated proposed practices in areas they felt qualified to assess, with structured feedback between rounds.
Results synthesized under preregistered constraints, including subgroup differences — the consensus practices above.
Stakeholders are invited to review and endorse the resulting guidance, strengthening shared expectations across the ecosystem.
Read the preregistered protocol · Full process & study design
The detailed data required for reanalysis of the results is publicly available: download the replication data. No replication code is currently available; we hope to address this before finalizing the written paper.
The promised deviations log records every change between the registered protocol and the implemented study.
Organizations that have endorsed the consensus statement.
Researchers, practitioners, auditors, policymakers, and other experts who have endorsed the consensus statement.
Delphi panel participants represented a cross-sector group of organizations.
Partial list of affiliations represented during the Delphi process. Participation does not imply endorsement of the final results.