A cross-sector initiative to establish community-endorsed guidance for conducting and reporting high-quality AI evaluations.
The Delphi rounds are complete and results are being finalized.
AI evaluations increasingly inform safety, governance, and deployment decisions—yet current practices remain uneven, siloed, and difficult to compare. We organized a broad, cross-sector Delphi panel to establish community-endorsed, concise guidance around practices for conducting and reporting high-quality evaluations, and are now finalizing the results. Such broadly endorsed guidance will create a shared reference and hopefully improve evaluation quality across labs, auditors, researchers, ethics teams, downstream model users, and practitioners. For more details, see our published call to join the process in Patterns.
This study is based on a consortium model, with organizational support and input from the organizations endorsing the project, and with contributions from many advisors — especially the authors of the protocol, Daniel Stuart Schiff, Michael Noetel, Stephen Casper, and David Manheim; the consortium manager, Irving Torres; and the study team, Andrea Loehr, David Manheim, Yixiong Hao, Drishti Sharma, Ishan Khire, Yilin Huang, Nidhi Sakpal, and Paulo Bova, among others.
AI Standards Lab
1) Organizational endorsement does not imply agreement with the final results; it is instead an endorsement of the importance and legitimacy of the process laid out in the protocol.
2) This is a partial list of affiliations represented during the Delphi process.
Evaluations are increasingly relied upon across labs, auditors, regulators, and researchers, but shared expectations about what "well-done" looks like remain thin.
We're not creating a new evaluation framework, benchmark, or metric—just making existing agreement visible.
Not a governance mandate, compliance regime, or audit protocol. Guidance, not requirements.
Not an attempt to resolve deep methodological disputes—focusing on areas of practical agreement.
A preregistered, multi-sector Delphi process to establish concise, community-endorsed guidance for conducting and reporting high-quality AI evaluations analogous to CONSORT, PRISMA, or TRIPOD.
A concise 1–2 page guideline form for auditors, labs, researchers, and other producers of AI evaluations to explain adopted or non-adopted practices, based on Delphi outcomes.
A structured, multi-round expert survey designed to identify where agreement already exists on concrete practices.
Broad review of evaluation frameworks, reporting checklists, measurement standards, and methodological critiques across ML, medicine, metrology, software testing, and experimental sciences.
Candidate items extracted from the systematic review, coded and mapped using a best-fit framework approach. Iteratively refined with expert feedback. Browse the Item Explorer.
Stakeholders across AI labs, auditors, academia, industry users, standards bodies, and government institutions contributed ratings and feedback to ensure broad representation in the consensus process.
Participants reviewed and rated proposed practices in areas they felt qualified to assess across the Delphi rounds.
Final synthesis is underway with preregistered constraints, transparent reporting of results including subgroup differences, and no post-hoc renegotiation.
Stakeholders will be invited to review and endorse the resulting consensus guidance, strengthening shared expectations for high-quality AI evaluations across the ecosystem.
AI evaluations provide important information for developers, deployers, and end users of AI systems. For this reason, it is important that they are built well and follow best practices. Therefore, based on the consensus of stakeholders and experts, we recommend the consideration and adoption of the below practices for constructing, documenting, maintaining, using, and reporting the results of evaluations.
Sign up to be alerted when the final results are available for review and endorsement.
Receive an alert when the final Delphi results and consensus materials are available.
Be notified when endorsement of the final outputs opens after synthesis and review are complete.
Help circulate and apply the final guidance once the report is released.
Receive a notice when organizations can review and endorse the final consensus outputs.
Use the alert to route the final materials to relevant policy, safety, governance, or evaluation teams.
Help share and maintain the checklist after publication.
The Delphi rounds are complete and results are being finalized. Reach out to learn more.
✉ [email protected]