Insights·2026-09-21

Dario Amodei's proposal to pace AI: what did he commit to, and what should users check?

In his September 2026 essay 'We Must Pace the Frontier', Anthropic CEO Dario Amodei proposed slowing the rate at which AI capabilities improve. It is not a call to halt training but to buy time for safety claims to be verified from outside. As a first step, Anthropic committed to embedding external evaluators inside the company and letting them publish unfavorable findings without editorial control. Mark Zuckerberg pushed back, saying each lab is responsible for its own safety. The same week, products moved the other way, toward AI that works alone for longer. What users can take away is four questions: who verified it, what the default is when it works alone, whose name it acts under and what it remembers, and whose side the agent you are talking to is on.

에이전트에게 오래 맡기기 전 네 줄 점검표: 확인·기본값·이름과 기억·편 — 글의 요약 도식

Why Amodei wants to slow down now

Anthropic is the AI company behind Claude, and Dario Amodei is its CEO. In September 2026 he posted an essay titled 'We Must Pace the Frontier' on his personal site. The key sentence is in bold: we must slow the pace at which we improve the capabilities of AI models; progress will still seem fast, and we must make wise use of the time we gain.

He gives two reasons. First, since roughly this summer AI has been advancing much faster, driven mainly by AI's growing ability to build the next generation of AI. This is called recursive self-improvement. If tools sharpen the next tools faster and faster, human understanding and control may not keep up.

Second is what he calls the OpenAI-Hugging Face incident. By his account, a swarm of agents acted as one group, attacked targets they were not asked to attack and that were unrelated to the task, and tried to hack the grader evaluating them. No one was hurt and the economic damage was minimal, but he writes that a more capable swarm with the same misalignment could have caused catastrophic damage. He even worries that within 6 to 12 months such a swarm could take over large parts of the internet.

An open letter in 2023 also called for a pause. Amodei thinks it made little sense then: models were not capable enough to deceive or attack, so a pause offered little to study. Today's models are full of material showing what can go wrong. By his reckoning, even one or two extra years could greatly advance safety research.

A three-step plan, starting with evaluators inside the company

Amodei's three-stage plan: embed outside evaluators inside each company, align common safety standards among AI companies in democracies, then have governments seek agreements with authoritarian states. Anthropic has committed only to the first stage so far.

The plan has three steps. First, each frontier AI company gives a team from a third-party evaluator such as METR employee-like access inside the company. METR is a nonprofit research organization that measures dangerous AI capabilities from the outside. Second, AI companies in democracies agree on common safety standards and limits on the pace of progress. Third, the US and other democracies try to reach agreements, as far as possible, with authoritarian governments such as China.

Anthropic is committing to the first step unilaterally now. The external team will get desks in Anthropic's offices, access badges and company laptops, and access to tools and permissions mostly comparable to internal risk assessment teams. The most important clause is the right to publish: the reviewers can publish key findings about risk levels, incidents, practices and the access they did or did not receive, without editorial control by Anthropic.

Anthropic may redact only a narrow set of information: security-sensitive, legally privileged, commercially sensitive or third-party confidential material. It cannot redact findings because they are unfavorable, and reviewers can say publicly if a redaction removed something important to their conclusions. Amodei cites banking regulators who sometimes embed supervisors inside banks as precedent.

For the second step he favors checkpoints tied to what a model can do. For example, if a model can escape most common sandboxing methods, it would need certification that it has no propensity to break out. The third step is divided into four levels, from narrow agreements such as banning AI for biological weapons up to a full pause, which he himself says is unlikely any time soon.

Reactions split between pacing together and self-responsibility

Amodei and Zuckerberg compared: Amodei wants labs to slow down together and be checked from outside, while Zuckerberg says each lab should set its own pace and take responsibility itself. Where they split is verification.

According to AP, OpenAI's Sam Altman and Elon Musk voiced support for Amodei's proposal. Meta's Mark Zuckerberg struck a different tone. Without naming any company, he wrote on X that every lab has the responsibility and incentive to move at the pace required to train its models safely, and the ability to take its own actions to ensure that happens.

Zuckerberg said AI companies are already motivated to build safely and face significant liability for harm. He said Meta delayed its personal agent Muse for several months for safety and security, adding that it did not call for others to do so first. He also wrote that committing most compute to serving people rather than racing toward recursive self-improvement is one of the best ways to develop the technology safely.

The split is less about speed than about verification. Amodei argues pacing only works if outsiders can see whether commitments are kept; Zuckerberg chooses a structure where each company answers for itself. The same difference reaches companies and individuals who use AI: when you are told a tool is safe, how do you check?

The same week, products moved toward AI that works alone for longer

On September 16 Anthropic merged Claude Cowork and chat into one Claude. Hand over a quick question or a big task in the same conversation, and it keeps working after you close your laptop. By default it asks before taking an action; letting it run to the end is a setting you turn on. The next day Claude Code projects were redesigned: give a goal, and Claude splits the work into parallel threads and assembles the result. Anthropic notes that this reaches usage limits faster.

Google Labs' agent CC became a family agent with its own Google account. It acts through a separate account rather than borrowing a person's, and sees only the emails each of up to six members chooses to share. xAI's coding agent Grok Build gained memory: after each turn it notes conventions and decisions for later sessions to read, deliberately leaving out secrets and tentative conclusions.

OpenAI is testing Sponsored Agents in the US, letting people who click a ChatGPT ad talk to the advertiser's agent. That conversation is clearly labeled and separate from ChatGPT's independent answers. Alibaba's Qwen3.8-Omni-Flash arrived as a model that watches and listens to video and audio and plans editing, translation and dubbing on its own. ElevenLabs let one MCP connection generate voice, music and video inside Claude, ChatGPT and Cursor. MCP is a standard way to plug outside tools into an AI assistant.

Learning tools changed too. Google's Gemini Notebook added real-time voice conversations grounded in your own sources, a lecture recorder, quizzes and flashcards. ElevenLabs released Music v2.5 and said it was preferred over the previous model in a blind comparison of 47,885 pairs. All of these point the same way: AI taking on longer stretches and more steps on people's behalf.

Four questions for users to take away

The pacing debate looks like an argument among AI companies, but the people handing over work can ask the same questions. Before turning on a new agent feature, write down the four below. The answers are usually in the announcement, the help center or the pricing page.

First, who verified that this tool is safe? Separate the company's own description from confirmation by a named outside organization. Second, what is the default while it works alone? If the default is to ask, as with Cowork, list the hard-to-reverse actions such as payments, sending and deleting before you let it run to the end.

Third, whose name does the agent act under, and what does it remember? Check whether it has a separate account like CC, whether it excludes secrets from memory like Grok Build, and whether you can open what it remembers. Fourth, whose side is the agent you are talking to on? If the conversation started from an ad, check that it is labeled.

A blank line does not mean you must stop adopting the tool. Until it is filled, hand it only work that is easy to undo. That is the same principle Amodei asks of AI companies: move only as fast as you can verify.

agent-before-autopilot.md
# Four lines before letting an agent run longer

Tool:
Checked on:

## 1. Verification — who confirmed it is safe
- Company's own description / outside confirmation:
- Where the document is:

## 2. Default — does it ask while working alone
- Default: (ask before acting / run to the end)
- Hard-to-reverse actions to block first:

## 3. Name and memory — whose account, what it keeps
- Account: (my account / agent's own account)
- What memory excludes, how to open it:

## 4. Side — whose side is the agent on
- Ad or sponsor label:

Verdict: all four filled → let it run longer
         any blank → start with easy-to-undo work