Process Archive

Process and Study Design

How the Evals-Consensus.AI Delphi initiative developed consensus-based guidance for conducting and reporting high-quality AI evaluations.

For the consensus statement and endorsement list, return to the front page.

Summary

AI evaluations increasingly inform safety, governance, and deployment decisions—yet current practices remain uneven, siloed, and difficult to compare. We organized a broad, cross-sector Delphi panel to establish community-endorsed, concise guidance around practices for conducting and reporting high-quality evaluations, and are now finalizing the results. Such broadly endorsed guidance will create a shared reference and hopefully improve evaluation quality across labs, auditors, researchers, ethics teams, downstream model users, and practitioners. For more details, see our published call to join the process in Patterns.

Consortium & Participants

This study is based on a consortium model, with organizational support and input from the organizations endorsing the project, and with contributions from many advisors — especially the authors of the protocol, Daniel Stuart Schiff, Michael Noetel, Stephen Casper, and David Manheim; the consortium manager, Irving Torres; and the study team, Andrea Loehr, David Manheim, Yixiong Hao, Drishti Sharma, Ishan Khire, Yilin Huang, Nidhi Sakpal, and Paulo Bova, among others.

Participating Individuals From2

  • AVERI
  • Centre for Governance of AI
  • Center for Open Science
  • Center for Security and Emerging Technology
  • Carnegie Mellon University
  • ETH Zurich
  • Imperial College London
  • Google DeepMind
  • Harvard University
  • Hugging Face
  • IBM Research
  • Intel
  • METR
  • MIT CSAIL
  • OpenAI
  • Oxford Internet Institute
  • RAND Center on AI, Security, and Technology
  • Schmidt Sciences
  • Tsinghua University
  • UC Berkeley
  • NYU
  • AI Security Institute
  • European AI Office

1) Organizational endorsement does not imply agreement with the final results; it is instead an endorsement of the importance and legitimacy of the process laid out in the protocol.
2) This is a partial list of affiliations represented during the Delphi process.

Why This Matters Now

Evaluations are increasingly relied upon across labs, auditors, regulators, and researchers, but shared expectations about what "well-done" looks like remain thin.

  • Many evaluations lack documentation on design choices
  • Statistical soundness and uncertainty often underspecified
  • Inconsistent scoring and applicability criteria
  • Results difficult to compare across organizations
  • Missing basics lead to avoidable failures
📋

Not a Framework

We're not creating a new evaluation framework, benchmark, or metric—just making existing agreement visible.

⚖️

Not a Mandate

Not a governance mandate, compliance regime, or audit protocol. Guidance, not requirements.

🤝

Building Consensus

Not an attempt to resolve deep methodological disputes—focusing on areas of practical agreement.

What We're Building

A preregistered, multi-sector Delphi process to establish concise, community-endorsed guidance for conducting and reporting high-quality AI evaluations analogous to CONSORT, PRISMA, or TRIPOD.

The Process

A structured, multi-round expert survey designed to identify where agreement already exists on concrete practices.

Systematic Review (Completed)

Broad review of evaluation frameworks, reporting checklists, measurement standards, and methodological critiques across ML, medicine, metrology, software testing, and experimental sciences.

Item Development(Completed)

Candidate items extracted from the systematic review, coded and mapped using a best-fit framework approach. Iteratively refined with expert feedback. Browse the Item Explorer.

Multi-Sector Participation (Completed)

Stakeholders across AI labs, auditors, academia, industry users, standards bodies, and government institutions contributed ratings and feedback to ensure broad representation in the consensus process.

Consensus Delphi Process (Completed)

Participants reviewed and rated proposed practices in areas they felt qualified to assess across the Delphi rounds.

Consensus Review (Finalizing)

Final synthesis is underway with preregistered constraints, transparent reporting of results including subgroup differences, and no post-hoc renegotiation.

Launch and Build Support (Completed)

Stakeholders were invited to review and endorse the resulting consensus guidance, strengthening shared expectations for high-quality AI evaluations across the ecosystem.

Report Writing and Finalization of Next Steps (In Process)

Writing the Delphi study paper and publishing detailed outcome and sensitivity appendices documenting the full results and subgroup analyses, alongside finalizing next steps for maintaining and building on the consensus guidance.

The Consensus Statement

AI evaluations provide important information for developers, deployers, and end users of AI systems. For this reason, it is important that they are built well and follow best practices. Therefore, based on the consensus of stakeholders and experts, we recommend we recommend evaluation practitioners consider adopting at least these shared practices for constructing, documenting, maintaining, using, and reporting the results of evaluations.

Join the Conversation

The Delphi rounds are complete and results are being finalized. Reach out to learn more.

[email protected]