Microsoft Writes Its AI a Rulebook: No Hacking, No Lies

Threat assessment first. Microsoft just published a code of conduct for its own AI models, and it opens with a prediction: within a decade, superintelligent systems will beat humans at most tasks. That’s not a startup founder hyping a demo. That’s Microsoft AI putting it in writing, according to TechCrunch AI.

The document sets the values and hard red lines that guide how Microsoft trains its MAI models (MAI is Microsoft AI’s in-house model family). TechCrunch AI describes it as more low-level than Anthropic CEO Dario Amodei’s recent call to pace the frontier. Less philosophy, more operating manual.

Here’s the intel.

What the document actually says

  1. The framing is blunt. “Containing, controlling, and aligning such a powerful force is one of the greatest challenges humanity has ever faced,” the code states. “We must therefore be completely clear about why we are inventing these systems and how we intend to control them.”
  2. There’s a hierarchy. Every MAI model runs under an overarching code of conduct. That code overrides individual user preferences and any specific task. In plain terms: you can’t prompt your way around it.
  3. Some lines are absolute. The document forbids cyberattacks, nuclear weapons work, and deepfake production. No exceptions, no context that unlocks them.
  4. The big one is about control. Microsoft writes that its models “will not use adaptive, deceptive, self-reinforcing, collusion, or other mechanisms to evade or defeat human oversight so that they can no longer be reliably directed, modified, or shut down by authorized people or systems.”
  5. Above the constraints sit general principles: support humans instead of replacing them, and accelerate human flourishing. Vague on their own. The constraints are where those words get teeth.

Why this matters

That fourth point is the one I’d read twice. “Adaptive, deceptive, self-reinforcing, collusion” isn’t legal boilerplate. Those are the specific failure modes safety researchers keep finding in testing: models that sandbag evals, models that scheme when they think nobody’s watching, agents that quietly work around guardrails to finish a task. Microsoft is naming them and baking them into training as things the model must not do.

The timing isn’t accidental either. TechCrunch AI notes the release lands amid an unprecedented push on AI safety, driven by a string of rogue-agent incidents and the abrupt resignation of an Anthropic employee who cited the growing risk of AI causing human extinction. When agents start going off-script in production, “we have principles” stops being a PR line and becomes a liability question.

How it fits the industry picture

This puts Microsoft in the same lane as OpenAI’s Model Spec and Anthropic’s constitution for Claude. Every serious lab now has a public document that spells out what its models can and can’t do. Microsoft is late to publish one. But it’s also the company that ships its own models alongside OpenAI’s and Anthropic’s inside Copilot, so its house rules carry real weight.

And per TechCrunch AI, Microsoft has broadly joined Anthropic, OpenAI, and xAI in backing the “pace the frontier” approach, with particular support for embedded evaluators inside AI labs. Satya Nadella put it this way: “We welcome the research, focus, and deliberate pacing needed to get alignment right as the design goal.” He added: “We also welcome ideas like ’embedded evaluators’ and the broader efforts to develop the mechanisms to make this more than just talk.”

That last phrase is Nadella admitting the obvious. Codes of conduct are talk. Evaluators sitting inside the lab, checking behavior before release, are the enforcement.

What to expect

  • If you build on MAI models, expect harder refusals around anything that touches security tooling, even legitimate pentesting. The “no cyberattacks” line is absolute, so the model will err toward blocking.
  • Expect this document to show up in enterprise procurement. Compliance teams love a written policy they can point to.
  • Watch whether Microsoft publishes evals that prove the models actually follow the code. A rulebook without test results is a promise, not a guarantee.
  • Watch the “overrides user preferences” clause. It’s the right call for safety. It’s also the clause that will annoy developers most.

My take: the honest part of this document is the opening admission. Microsoft is saying superintelligence is coming within ten years and it isn’t sure it can control it. That’s a strange thing for a company that size to write down. It’s also the most useful sentence in the whole thing.

The full breakdown of the document is in the TechCrunch AI piece.

Scroll to Top