How the Evals-Consensus.AI Delphi initiative developed consensus-based guidance for conducting and reporting high-quality AI evaluations.
For the consensus statement and endorsement list, return to the front page.
AI evaluations increasingly inform safety, governance, and deployment decisions—yet current practices remain uneven, siloed, and difficult to compare. We organized a broad, cross-sector Delphi panel to establish community-endorsed, concise guidance around practices for conducting and reporting high-quality evaluations, and are now finalizing the results. Such broadly endorsed guidance will create a shared reference and hopefully improve evaluation quality across labs, auditors, researchers, ethics teams, downstream model users, and practitioners. For more details, see our published call to join the process in Patterns.
This study is based on a consortium model, with organizational support and input from the organizations endorsing the project, and with contributions from many advisors — especially the authors of the protocol, Daniel Stuart Schiff, Michael Noetel, Stephen Casper, and David Manheim; the consortium manager, Irving Torres; and the study team, Andrea Loehr, David Manheim, Yixiong Hao, Drishti Sharma, Ishan Khire, Yilin Huang, Nidhi Sakpal, and Paulo Bova, among others.
AI Standards Lab
1) Organizational endorsement does not imply agreement with the final results; it is instead an endorsement of the importance and legitimacy of the process laid out in the protocol.
2) This is a partial list of affiliations represented during the Delphi process.
Evaluations are increasingly relied upon across labs, auditors, regulators, and researchers, but shared expectations about what "well-done" looks like remain thin.
We're not creating a new evaluation framework, benchmark, or metric—just making existing agreement visible.
Not a governance mandate, compliance regime, or audit protocol. Guidance, not requirements.
Not an attempt to resolve deep methodological disputes—focusing on areas of practical agreement.
A preregistered, multi-sector Delphi process to establish concise, community-endorsed guidance for conducting and reporting high-quality AI evaluations analogous to CONSORT, PRISMA, or TRIPOD.
A concise 1–2 page guideline form for auditors, labs, researchers, and other producers of AI evaluations to explain adopted or non-adopted practices, based on Delphi outcomes.
A structured, multi-round expert survey designed to identify where agreement already exists on concrete practices.
Broad review of evaluation frameworks, reporting checklists, measurement standards, and methodological critiques across ML, medicine, metrology, software testing, and experimental sciences.
Candidate items extracted from the systematic review, coded and mapped using a best-fit framework approach. Iteratively refined with expert feedback. Browse the Item Explorer.
Stakeholders across AI labs, auditors, academia, industry users, standards bodies, and government institutions contributed ratings and feedback to ensure broad representation in the consensus process.
Participants reviewed and rated proposed practices in areas they felt qualified to assess across the Delphi rounds.
Final synthesis is underway with preregistered constraints, transparent reporting of results including subgroup differences, and no post-hoc renegotiation.
Stakeholders were invited to review and endorse the resulting consensus guidance, strengthening shared expectations for high-quality AI evaluations across the ecosystem.
Writing the Delphi study paper and publishing detailed outcome and sensitivity appendices documenting the full results and subgroup analyses, alongside finalizing next steps for maintaining and building on the consensus guidance.
AI evaluations provide important information for developers, deployers, and end users of AI systems. For this reason, it is important that they are built well and follow best practices. Therefore, based on the consensus of stakeholders and experts, we recommend we recommend evaluation practitioners consider adopting at least these shared practices for constructing, documenting, maintaining, using, and reporting the results of evaluations.
The Delphi rounds are complete and results are being finalized. Reach out to learn more.
✉ [email protected]