Embedded External Evaluators: Anthropic's Proposal to Put Safety Watchdogs Inside AI Labs

A September 2026 Anthropic proposal to give outside safety researchers employee-level access inside frontier AI labs — badges, desks, and the right to publish findings without company editorial control.

Created 2026-09-17 Last reviewed 2026-09-17

What it is

“Embedded external evaluators” refers to a proposal, put forward by Anthropic CEO Dario Amodei in a September 12, 2026 essay titled “We Must Pace the Frontier,” to give outside safety researchers ongoing, employee-like access inside frontier AI companies. Rather than the current model — where outside evaluators are typically brought in for short, bounded testing windows before a model’s release — embedded evaluators would work from inside the lab on a continuous basis, with desks, security badges, company laptops, and access to internal tools comparable to an internal risk-assessment team.

The proposal goes beyond physical presence. Embedded evaluators would be able to inspect intermediate training checkpoints (not just finished, released models), review post-training environments and evaluation logs, and interview employees about safety practices. Crucially, Amodei’s plan states that evaluators could publish their findings about risk levels, incidents, and lab practices without the company’s editorial control. The company could redact only material that is security-sensitive, legally privileged, commercially sensitive, or covers third-party confidential information — not conclusions the lab simply finds unflattering. Evaluators would also be permitted to disclose publicly if such redactions occurred.

Anthropic named METR (a nonprofit AI evaluation organization) as an example partner for this arrangement. Embedded evaluators are the first of three planks in Amodei’s broader proposal, alongside coordination among AI companies in democratic countries on safety standards, and — eventually — an international agreement that would also involve authoritarian governments, including China.

Why it matters for AI governance and narratives

The embedded-evaluator proposal sits at the center of a long-running framing contest over who gets to verify claims that AI labs make about their own systems’ safety. For years, labs have self-reported model capabilities and risks, with occasional third-party red-teaming conducted under contractor-style arrangements — typically bound by non-disclosure agreements and subject to the lab’s control over what gets published. Embedded evaluation is Anthropic’s answer to a criticism that has dogged the industry: that voluntary safety commitments are unverifiable promises made by companies grading their own homework.

But the proposal also illustrates a structural tension that outside researchers were quick to flag. The AI company chooses who receives embedded access, defines the boundaries of that access, and retains ultimate control over what evaluators can see and what counts as legitimately redactable. That means the same actor whose behavior is being scrutinized also designs the scrutiny. Whether this arrangement produces genuine independent oversight or a more sophisticated form of reputational management is now an open contest of framing — one where builders, evaluators, and regulators each have an incentive to describe the same mechanism differently.

Key facts and dates

Amodei published “We Must Pace the Frontier” on September 12, 2026, framing the embedded-evaluator commitment as something Anthropic would adopt unilaterally rather than wait for industry-wide agreement. OpenAI CEO Sam Altman said OpenAI would also commit to the practice, a notable alignment between two competing labs. As of mid-September 2026, Meta, Google DeepMind, and xAI had not committed to the embedded-evaluator model.

Organizations named or discussed as candidates for this kind of access include METR, Redwood Research, Apollo Research, FAR.AI, Palisade Research, and Safer AI. Researchers who spoke to TechCrunch broadly welcomed the direction of the proposal but raised specific concerns grounded in past experience. Adam Gleave, CEO of FAR.AI, noted that evaluators have historically operated as contractors under restrictive NDAs with the developer controlling what could be published. Apollo Research’s Alexander Meinke argued the underlying problem persists regardless of the new arrangement: “we are completely relying on AI companies to both carefully check this themselves and then truthfully report this to the public. And we’ve seen from recent incidents that, by default, they will do neither.” Henry Papadatos of Safer AI argued that voluntary measures remain contingent on a company’s goodwill and that binding legislation, not voluntary commitments, would be needed to make the arrangement durable. Researchers have also pointed to time-constrained past evaluations — for example, Apollo Research reportedly had only a few days to test one recent model — as evidence that access duration and scope, not just physical presence, determine whether an evaluation is meaningful.

Where to learn more

Sources

Primary source: Amodei's own essay proposing the embedded-evaluator commitment, including specific access and publishing terms.
Detailed reporting with named quotes from evaluator organizations on independence concerns and historical precedent.
Independent critical analysis of the proposal from a major news outlet.
Corroborating coverage of industry adoption and named evaluator organizations.
Referenced in: Editorial No. 325