Public Letter

Minimum Conditions for Embedding Evaluators

Published September 18, 2026100+ Signatories

We, the undersigned, are encouraged to see frontier AI companies call for embedding third-party organizations1,2,3,4,5 to evaluate rapidly escalating AI capabilities and risks. We believe that all frontier AI companies should embed evaluators to independently assess AI risks, including evaluating the systems themselves and any significant incidents of real-world harm, as well as the companies’ training, deployment, oversight, operational, and safeguard practices.

To be credible, embedded third-party evaluations must have scientific objectivity, transparency, independence, and robust protections against interference from the evaluated companies, including at least:

  1. Frontier AI companies should rely on evaluators that are meaningfully independent, that maintain full editorial control, and that disclose and mitigate potential conflicts of interest. This includes at a minimum that embedded evaluation organizations should not be owned or governed by frontier AI companies, should not have other significant commercial business with them, and should not accept any form of payment or other reward contingent on the evaluator’s findings.
  2. Frontier AI companies should incorporate differing viewpoints and areas of expertise, including by embedding multiple evaluation organizations across a range of priority risk areas, each with deep relevant technical expertise, as well as by allowing and encouraging evaluators to share how conclusions differ among evaluators and between evaluators and company employees.
  3. Embedded evaluators should be transparent, including transparency about their methods and findings, the nature of their access, and the broader terms of the evaluation. Frontier AI companies should actively facilitate this transparency, including limiting the scope of non-disclosure agreements. They should also allow evaluators prompt and unfiltered communication with the companies’ boards and other privileged oversight bodies, as well as public release of findings and evidence, subject only to a time-limited redaction process restricted to protecting critical interests in intellectual property, customers’ sensitive information, individual privacy, security, and public safety.
  4. Embedded evaluators should be shielded from retaliation from the companies they embed with for choosing reasonable evaluation methods, discovering information, or drawing conclusions that are unflattering to those companies. This includes reasonable protections against retaliatory litigation, as well as funding mechanisms that give them confidence they will remain funded even in these cases.
  5. Frontier AI companies should grant embedded evaluators access equivalent to that of their own highly privileged employees for the purposes of their evaluations, and with exceptions to protect sensitive data belonging to the company’s customers and other third parties. This includes access to the same relevant systems, data, tools, and physical spaces as those available to senior internal company employees responsible for carrying out comparable risk assessments, as well as candid and direct one-on-one communication with relevant staff.

This list is not comprehensive, and conditions like these to ensure credible evaluations should be increasingly standardized, codified, and enforced. One example is the set of terms defined in the AEF-1 standard, which has already seen early adoption, but far more work will be necessary to ensure that embedded evaluators are effective and meaningful.

Embedded evaluations cannot address all oversight needs and should be treated as a complement to, rather than a replacement for, broader efforts by frontier AI companies to expand external oversight, including greater public transparency and additional, broader forms of access for independent researchers.

Notable Signatories

All signatories sign in their personal capacity. Affiliations are listed for identification only.

Geoffrey Hinton

Nobel Laureate in Physics (2024); University Professor Emeritus, University of Toronto

Stuart Russell

Distinguished Professor of Computer Science, University of California, Berkeley

Arvind Narayanan

Professor of Computer Science, Princeton University

Yejin Choi

Professor, Stanford University

Vinh Nguyen

Former Chief AI Officer, National Security Agency

Jacob Steinhardt

CEO, Transluce

Daniel E. Ho

Professor, Stanford University

Daniel Kokotajlo

Exec Director, AI Futures Project

Ben Buchanan

Professor, Johns Hopkins University

Andrew Freedman

CEO, Fathom

Rumman Chowdhury

CEO, Humane Intelligence PBC

Jose H. Orallo

Professor, University of Cambridge

Rob Reich

Professor, Stanford University

Seth Lazar

Professor, Johns Hopkins University

Rayan Krishnan

CEO, Vals AI

Micah Hill-Smith

Co-Founder & CEO, Artificial Analysis

Stella Biderman

Executive Director, EleutherAI

Robert Trager

Director, Oxford AI Governance Initiative

Miles Brundage

Executive Director, AI Verification and Evaluation Research Institute (AVERI)

Timothy Fist

President, Institute for Progress

Joal Stein

Executive Director, Collective Intelligence Project

Rajiv Dattani

Co-founder, Artificial Intelligence Underwriting Company (AIUC)

Charles Teague

CEO, Meridian Labs

Adam Gleave

Co-founder & CEO, FAR.AI

Mathilde Collin

Co-founder and CEO, KORA

Nathan Lambert

Executive Director, Trillium Labs

Henry Papadatos

Executive Director, SaferAI

Trooper Sanders

President, State AI Safety Roundtable

Alex Kleinman

Co-founder, Active Site

Owain Evans

Director and Lead Researcher, Truthful AI

Page Hedley

CEO, Guidelight AI Standards

Dr. Zachary Stein

Founder, AI Psychological Research Coalition

Steven Wolfe Pereira

Chief Executive Officer, Alpha Governance PBC

Sarah Schwettmann

Chief Scientist, Transluce

Conrad Stosz

Head of Governance, Transluce

Miranda Bogen

Chief Technologist, Center for Democracy & Technology

Dr. Rebecca Portnoff

Personal Capacity

Andrew Strait

Personal Capacity

Nick Beckstead

CEO, Secure AI Project

Thomas Woodside

Co-Founder, Secure AI Project

Mitch Prinstein

John Van Seters Distinguished Professor and Co-Director, Winston Center on Technology and Brain Development, University of North Carolina at Chapel Hill

Samira Nedungadi

Head of Engineering, AI Team, SecureBio

Markus Anderljung

Director of Policy & Research, GovAI

Hoda Heidari

Assistant Professor of Machine Learning and Societal Computing, Carnegie Mellon University

Jeremie Harris

Co-founder and CEO, Gladstone AI

Charles Foster

Member of Policy Staff, METR

Remzi Arpaci-Dusseau

Founding Dean, College of Computing & Artificial Intelligence, University of Wisconsin-Madison

Kevin Frazier

Professor, The University of Texas School of Law

Sophia Hatz

Associate Professor, Alva Myrdal Center for Nuclear Disarmament, Uppsala University

Jacob Andreas

Associate Professor, Electrical Engineering and Computer Science, Massachusetts Institute of Technology

Christian Schroeder de Witt

Associate Professor, University of Oxford

David Duvenaud

Associate Professor of Computer Science, University of Toronto

Dylan Hadfield-Menell

Associate Professor of Electrical Engineering and Computer Science, Massachusetts Institute of Technology

Dr Karine Caunes

Executive Director, Digihumanism - Centre for AI & Digital Humanism

Stephen Casper

Assistant Professor of Public Policy, Harvard University

Umang Bhatt

Assistant Professor in Trustworthy Artificial Intelligence, University of Cambridge

Florian Tramèr

Assistant Professor of Computer Science, ETH Zurich

Cameron Jones

Assistant Professor of Psychology, Stony Brook University

Marc Aidinoff

Professor, Harvard University

Nicholas Caputo

Assistant Professor, Johns Hopkins School of Government and Policy

Sayash Kapoor

Incoming professor, UC Berkeley

Anka Reuel

Computer Science PhD Candidate, Stanford University

Nada Madkour

Director of the AI Security Initiative, UC Berkeley Center for Long-Term Cybersecurity

Ranjit Singh

Director, AI on the Ground Program, Data & Society Research Institute

Rob Krzyzanowski

Executive Director, Poseidon Research

Francis Rhys Ward

Director of Arrow Research, Arrow Research

Shea Brown

Founder & CEO, BABL AI

Jaime Raldua Veuthey

CEO, Apart Research

Alex Engler

Executive Director, Penn MEDIATED (Center on Media, Technology, and Democracy)

Lewis Hammond

Research Director, Cooperative AI Foundation

Avijit Ghosh

Lead/Founding Member, Evaluating Evaluations (EvalEval) Coalition

Charbel-Raphael Segerie

Executive Director, CeSIA - French Center for AI Safety

Alexander Meinke

Head of Research, Apollo Research

Jasper Götting

Director of AI, SecureBio

Ross Matican

Investor and Grantmaker, Halcyon Futures

Nicolai Ouporov

CEO, Fleet AI

Piercosma Bisconti

Managing Director, Icaro Foundation

Darius Emrani

CEO, Scorecard

Nicolas Moës

Executive Director, The Future Society

Imran Khan

Program Director, Psychosocial Evaluations of AI, Center for Humane Technology

Jared Moore

PhD Student in Computer Science, Stanford University

Rosie Campbell

Managing Director, Eleos AI Research

Patricia Paskov

Director of Standards, AI Verification and Evaluation Research Institute (AVERI)

Koen Holtman

Co-lead, AI Standards Lab

Ze Shen Chin

Co-lead, AI Standards Lab

Sophie Williams

Research Fellow, GovAI

Peter Slattery

Personal Capacity

Dan Bateyko

Researcher, Cornell University Department of Information Science

Jacob Hilton

Personal Capacity

Ben Bucknall

PhD Student, University of Oxford

Malcolm Murray

Research Lead, SaferAI

Pamela Mishkin

Personal Capacity

Jimmy Farrell

EU AI Policy Lead, Pour Demain

Briana Vecchione

Technical Researcher, Data & Society Research Institute

Jerome Wynne

Research Scientist, Apart Research

Girish Sastry

Member of Research Staff, Guidelight AI Standards

Markov Grey

Head of Technical AI Governance, French Center for AI Safety (CeSIA)

Alan Chan

Research Fellow, GovAI

Marta Bieńkiewicz

Policy and Partnership Manager, Cooperative AI Foundation

Ichhya Pant

Personal Capacity

Katarina Slama

Personal Capacity

Wout Schellaert

Personal Capacity

Noah Ringler

Personal Capacity

Daniel Filan

Member of Research Staff, Guidelight AI Standards

Joe Kwon

Member of Research Staff, Guidelight AI Standards

Lucia Velasco

Personal Capacity

Gabriel Levie

Personal Capacity

Ian Eisenberg

Personal Capacity

Jacob Charnock

Personal Capacity

Shen Zhou Hong

Personal Capacity

Florian Brand

Research Engineer, Prime Intellect

Sign the Letter

By signing, you agree to have your name and title displayed publicly as a supporter. We will not share your personal information with third parties without your consent.