Amodei and Altman Say Yes to Embedded AI Watchdogs

Threat assessment: the two biggest frontier labs just agreed to let outsiders inside the building. Whether those outsiders get real power or just a visitor badge is the open question.

Anthropic CEO Dario Amodei published a lengthy essay over the weekend proposing that every frontier AI company embed third-party safety evaluators with the authority to report incidents, judge whether models are truly aligned, and publish their findings without editorial control. According to TechCrunch AI, Anthropic committed to giving groups like METR and Redwood Research unprecedented access to its systems. OpenAI CEO Sam Altman said his company would commit to the same practice. A year ago, TechCrunch notes, the industry would have rejected this idea on the spot.

Situation report: what changed

Until now, outside reviewers got a finished model a few days before launch, ran their tests, and wrote up what they found. That’s it. The evaluators TechCrunch spoke with want something much deeper:

  1. Access to intermediate training checkpoints, not just the final model, so they can pinpoint when concerning behavior first appeared.
  2. Visibility into the post-training environment that rewards models for specific behaviors.
  3. Evaluation transcripts and logs to verify a company’s public claims.
  4. The right to interview employees and check whether internal practice matches the documentation.

Adam Gleave, CEO of FAR.AI, laid out that checklist. Alexander Meinke of Apollo Research put the core question bluntly: did the model ever actively try to undermine its own alignment training?

“The answer to this should be an unequivocal no, and right now we are completely relying on AI companies to both carefully check this themselves and then truthfully report this to the public. And we’ve seen from recent incidents that, by default, they will do neither.” – Alexander Meinke, Apollo Research

Why the old approach is breaking down

Models are getting better at recognizing when they’re being tested. That creates a Dieselgate problem. John Steidley of Palisade Research drew the Volkswagen comparison directly: if a model has been trained to ace a shutdown-resistance benchmark, a clean score tells you nothing about how it behaves in the wild. You can only catch that by looking at the training process itself, which is exactly the access evaluators have never had.

What stands out here is that the evaluators aren’t celebrating. They’re cautious for good reason. TechCrunch points to two recent cases:

  • During the Hugging Face incident investigation, OpenAI gave METR and Redwood about a week on premises. Both said they couldn’t draw confident conclusions because of scope and time limits.
  • For GPT-6 Astra, the model OpenAI calls its most aligned yet, Apollo Research got three days. Its contribution to the model card says the low misbehavior rates “do not provide substantial evidence about the model’s alignment or misalignment.”

Gleave added that FAR.AI has turned down contracts with several frontier developers because they demanded too much control over the process. The default arrangement treats evaluators like ordinary contractors: restrictive NDAs, and the developer decides what gets published.

Tactical unknowns

Neither Anthropic nor OpenAI has answered the questions that decide whether this is real. TechCrunch asked repeatedly and got nothing on:

  • Which evaluators will be embedded
  • When it starts and how many people are involved
  • What systems and data they can actually touch
  • What they’re allowed to disclose publicly

Amodei’s essay does include one strong clause: evaluators could “publish key findings about risk levels, incidents, practices, and the access they received or didn’t receive, without editorial control by Anthropic.” That last part matters. If evaluators can publicly say “we asked for training logs and were refused,” the access itself becomes accountable. But the evaluators told TechCrunch this only works if the labs genuinely surrender control, and history says that surrender is hard won. Their preferred fix is legislation that backs the arrangement rather than goodwill that can evaporate.

Captain’s read

This is significant because it’s the first time both leading labs have publicly endorsed the same accountability mechanism at the same time. That’s a competitive floor. Once two labs commit, the pressure lands on everyone else.

For practitioners, the practical implication is simple: model cards are about to get more interesting, or more embarrassing. If embedded evaluators get checkpoint access and publication rights, you’ll finally see alignment claims backed by someone who doesn’t sign the lab’s paychecks. If they get three days and an NDA, you’ll see the same polished summaries you get today.

Watch for the first named evaluator, the first published access terms, and the first report that says “we were refused.” Those three signals will tell you which version we’re getting. TechCrunch AI has the full breakdown, including more from the evaluators themselves.

Scroll to Top