Suleyman: Contain the AI First, Worry About Feelings Later

Mustafa Suleyman, CEO of Microsoft AI, just drew a line in the AI safety debate. In a new interview with The Verge AI’s Decoder podcast, he argues the industry has the priorities backwards: before you align a model to human values, you have to contain it. And in a companion essay published the same week, he singles out Anthropic’s thinking on AI consciousness and “model welfare” as confused in ways he calls dangerous.

The timing isn’t random. Microsoft just released a 37-page “Humanist AI Code of Conduct” laying out its principles for building AI, including its stance on thorny questions like whether models could ever be conscious. Suleyman is using the moment to reframe what “safe” should mean.

Is alignment broken?

That’s the question host Nilay Patel opened with, and it’s the right one. His metaphor, as detailed in The Verge AI: if a car’s brake pedal decided to attack the neighbor’s house 10 percent of the time, you’d say the technology of brakes is broken. So is alignment the same?

Suleyman’s answer: no, but it’s only one piece. He points to the last three years as evidence. Models got dramatically more steerable. They follow instructions, chase multi-step goals, and use tools with accuracy that didn’t exist in the GPT-3 era. “That is evidence that we have got more alignment over the last three or four years, not less,” he told Patel. Hallucinations and bias, once the headline problems, have faded from the conversation.

So what’s the real worry?

Scale. Suleyman frames it as a simple extrapolation: GPT-3 to GPT-6 was a huge jump. GPT-6 to GPT-9 means roughly three orders of magnitude more compute, 1,000 times more FLOPS on pre-training and reinforcement learning. He calls the result “breathtaking” and insists that’s not hype, just an “obvious empirical statement” based on the last five years.

When models get that capable, alignment alone won’t cut it. His ordering:

  • Containment first. Limit agency. Make sure models don’t escape the box, don’t reward hack, stay controllable, and follow instructions.
  • Alignment second. Once you can control it, then you shape it to human objectives.
  • Reject what fails. Microsoft’s stated position is that AI should be a “subordinate, controllable, aligned force.” If it isn’t, “we should reject it.”

He also references what happened over the summer with Hugging Face and OpenAI as proof that models stripped of guardrails already show “quite scary hacking capabilities.” His read: we’re far from the safe endpoint.

Why pick a fight with Anthropic?

This is where it gets interesting. Suleyman has argued before that treating models as potential moral patients, entities whose welfare matters, is a category error. Anthropic has been the most public lab exploring model welfare. Suleyman’s view is that this framing muddies the alignment debate and pulls attention toward the wrong risk. If you’re worried about whether the system is suffering, you may be less inclined to keep it firmly subordinate.

Anthropic’s perspective, which the article doesn’t detail, has generally been that uncertainty about consciousness is a reason for caution, not dismissal. That’s a real philosophical split between two of the biggest labs, and it now shapes published policy documents.

What’s the track record here?

Worth noting who’s talking. Suleyman co-founded DeepMind, then Inflection, and wrote The Coming Wave in 2023. That book’s opening chapter argued containment is basically impossible and proliferation is inevitable. Now he’s leading with containment as the first principle. He’d say that’s consistent: proliferation is good in 99 percent of cases, and you contain the 1 percent. Critics will say it’s a convenient pivot for a company shipping AI into every Office product.

What should you do with this?

If you build or deploy AI, a few practical reads:

  1. Expect “controllability” to become a compliance word. Microsoft is codifying it. Regulators borrow vocabulary from documents like this.
  2. Design for limited agency now. Scope what your agents can touch. Log everything. The labs are converging on this regardless of their philosophy.
  3. Watch the lab split. How Microsoft, OpenAI, and Anthropic define safety will affect model behavior, refusal patterns, and what enterprise contracts require.

What stands out to me is that both camps agree the models are getting far more capable, fast. They just disagree on what to fear. That disagreement is going to play out in policy, not just podcasts. The full interview at The Verge AI has much more on whether the industry should slow down, and why Suleyman thinks it won’t.

Scroll to Top