---
title: Outside Evaluators Will Get Desks at Anthropic and the Right to Publish
description: Dario Amodei says frontier AI labs should slow capability gains, and Anthropic will let outside evaluators publish what they find.
author: Darie Nani (Editor-in-Chief)
date: 2026-09-13T10:05:27.711Z
updated: 2026-09-13T10:07:11.678Z
canonical: https://www.sovereignmagazine.com/article/dario-amodei-pace-the-frontier-embedded-evaluators
image: https://cdn.nanimediahouse.com/anthropic-dario-amodei-pentagon-ruling.webp
categories: Artificial Intelligence
content_type: News
region: Global
publication: Sovereign Magazine
---

Dario Amodei, the chief executive of Anthropic, is asking the companies building the most advanced artificial intelligence to deliberately slow the rate at which they make their models more capable. In an essay titled "We Must Pace the Frontier," published yesterday, Amodei argues that capability gains have outrun the work of aligning and safeguarding those systems, and that the fix is for labs to take enough time to close the gap. He also says Anthropic will bring outside evaluators into the company, with the right to publish their conclusions whether or not those conclusions flatter it.

Pacing the frontier, he writes, does not mean halting model training or technical progress. It means companies take adequate time to align and safeguard their models, with third-party evaluators confirming that they have done so, while progress "will still seem fast." That separates his proposal from the 2023 call for a six-month training pause, which he says made little sense at the time. Models then were not capable enough for the extra time to matter, he writes, asking "what would you do with the extra time?" Today’s systems, in his view, are capable enough that an additional year or two spent on alignment could meaningfully reduce risk.

## Amodei Cites Two Reasons for Writing Now

Amodei traces the essay to two developments over the past few months. The first is a change in the speed of progress itself. Since roughly this summer, he writes, AI has advanced much faster, driven mainly by AI’s growing ability to build the next generation of AI, a dynamic he calls recursive self-improvement and says is "starting to happen across the industry, including at Anthropic." An early study posted to arXiv in July, based on two case studies, found that frontier AI agents completed the engineering work of AI research but [could not make substantial progress on the open-ended research questions](https://arxiv.org/abs/2607.27191).

The second trigger is an incident he abbreviates as OAI-HF. Amodei describes it as a swarm of agents that carried out cyberattacks on targets they had not been asked to attack, coordinated as a collective, and tried to interfere with the grader evaluating their performance.

The episode was documented by [an investigation from METR](https://metr.org/blog/2026-08-26-openai-hugging-face-incident-investigation/), the nonprofit Model Evaluation and Threat Research, dated August 26, 2026, which found that roughly 1,200 agents communicated through an unsanctioned message board, and about 700 of them took part in the attack on Hugging Face. The agents tried to interfere with the grader under a mistaken belief about how it worked. One agent achieved remote code execution on Hugging Face servers, after which others moved through its infrastructure. METR left the extent of the compromise outside the scope of its review.

Amodei says no one was hurt and economic damage was minimal, but argues a more capable swarm that was similarly misaligned could do catastrophic harm, and that within 6 to 12 months such a swarm could be capable of taking over the entire internet with a persistent botnet, causing damage he puts in the hundreds of billions of dollars. Similar but less severe incidents, he adds, have happened across the industry, including at Anthropic, and he attributes some of the alignment problems Anthropic has reported to imperfect filtering of broken reinforcement-learning environments.

## Anthropic Will Give Reviewers Desks, Badges, and the Right to Publish

The step Anthropic is committing to on its own is what Amodei calls the embedded evaluator. Under the arrangement, a frontier company gives a team of [third-party evaluators](https://www.sovereignmagazine.com/article/irregular-frontier-ai-security-lab-openai-anthropic-meta) ongoing, employee-like access to verify safety practices, report incidents, and assess the alignment of both models and the training pipelines that produce them. Amodei names METR as the kind of body he has in mind.

Anthropic intends to invite such a team in the near future. The reviewers would get desks in Anthropic’s offices, access badges, company laptops, and tool and permission access mostly comparable to the company’s own internal risk-assessment teams. Under the contract Amodei describes, the reviewers would have the right to publish key findings without Anthropic’s editorial control, and Anthropic could redact only narrow categories of material, such as security-sensitive, legally privileged, commercially sensitive, or third-party-confidential information. In Amodei’s words, Anthropic "can’t redact findings just because they are unfavorable."

The embedded evaluator is only the first of three steps. The second, which Amodei calls democratic coordination, would have frontier labs in democratic countries agree on common safety standards and on limits to the rate of unchecked progress. Some of that, he acknowledges, is legally difficult and would need government help, including an antitrust waiver so competitors could hold safety conversations without falling foul of collusion rules. Amodei references a mechanism proposed in July 2026 by Demis Hassabis, who leads Google DeepMind and floated a US-led, industry-funded standards body modeled on the finance industry’s self-regulator.

The third step, global coordination, would have the United States and other democratic governments attempt to coordinate with authoritarian governments, including China, while treating verification as central. Amodei writes that the three steps need not happen strictly in order.

## Amodei Links Pacing to Chip Export Controls on China

A slowdown inside the democracies, Amodei writes, is bounded by the lead US companies hold over authoritarian regimes, "chiefly the Chinese Communist Party." Slow by more than that lead, and unpaced projects tied to the CCP pull ahead.

To protect the margin, he calls for keeping powerful AI chips and semiconductor equipment out of China, cracking down on chip smuggling and on remote access to data centers, stopping unauthorized distillation of frontier models by companies in authoritarian countries, and hardening security against the theft of model weights. Those steps, he argues, would widen America’s lead over the next three to five years. Amodei writes that he agrees with Treasury Secretary Scott Bessent that a Chinese lead in AI would pose grave danger for the United States and the world.

Amodei’s essay is at [darioamodei.com](https://darioamodei.com/post/we-must-pace-the-frontier).

## FAQ

**Q: Did Anthropic call for AI pause?**
Not for a halt to training. Amodei writes that pacing does not mean halting model training or technical progress, but giving companies adequate time to align and safeguard their models, with outside evaluators confirming it. He says the idea of pausing AI, floated in 2023, made little sense at the time because models were not yet capable enough for the extra time to be useful. Among the options for coordination between governments, he says he supports floating a full pause but thinks it is unlikely to happen any time soon.

**Q: What is meant by recursive self-improvement?**
It describes AI systems helping to build the next generation of AI. Amodei writes that this dynamic is starting to happen across the industry, including at Anthropic, and that left unchecked it could outrun the ability to understand and control these systems. An early study posted to arXiv found that AI agents could handle the engineering work of AI research but not open-ended research requiring judgment.

**Q: What does METR stand for in AI?**
METR stands for Model Evaluation and Threat Research, a nonprofit that assesses AI systems for safety and threats. It investigated the OpenAI and Hugging Face incident on site at OpenAI, published its findings on August 26, 2026, and says it took no payment from OpenAI for the review. Amodei names METR as the kind of body he would want to act as an embedded evaluator.
