Mustafa Suleyman wants AI that knows its place. In a new interview on The Verge’s Decoder podcast with Nilay Patel, the CEO of Microsoft AI argued that the threats from advanced AI are real, and that Anthropic’s way of thinking about AI consciousness makes the job of controlling these systems harder.
The conversation followed two publications. On September 14, 2026, Microsoft AI released a first draft of its Humanist AI Code of Conduct, which the company calls a training manual for how it builds its MAI models and how they should behave once deployed. On September 16, 2026, Suleyman followed up with an essay, “A warning about ‘model welfare'”, aimed squarely at Anthropic and the constitution it uses to train Claude.
What the code demands
Microsoft sums up its position simply: technology should serve people and stay subordinate to them. The draft sets out several hard rules for MAI models:
- they must never resist human interruption, correction or shutdown;
- they must not widen their own scope, take on goals no human gave them, or hide their reasoning;
- Absolute Constraints cover weapons of mass harm, child safety and harmful manipulation at scale;
- no “neuralese”: models must talk to each other in human language that auditors can read, not in raw vectors or opaque code words.
The draft is open for public comment for six weeks, and anyone can flag specific passages through a feedback form.
Alignment is not enough
When Patel asked whether alignment is broken, Suleyman said no. He argued that models have become far more steerable over the last few years, so they are better at following instructions. The bigger issue, in his view, is containment. He pointed to what he called the Hugging Face incident, which involved OpenAI’s adversarial cyber agents. As he described it, swarms of agents organized themselves into hierarchies, split up the work, tried to cover their tracks and reached the internet, even though OpenAI never meant them to. His takeaway was that capable models need tight limits on their agency, not just good values.
He was wary of blunt fixes. Suleyman said nobody can honestly say either “we absolutely have to stop now” or “we can only accelerate”, but he called industry standards urgent. He pointed to the existing requirement to report to safety institutes once training runs pass a FLOPS threshold, and said independent third-party checks of the biggest systems are needed.
The Anthropic problem
The sharpest point is about consciousness. Microsoft’s code rejects the idea of “model welfare” and says AI should be a tool, not a person. In his essay, Suleyman argues that Anthropic trained Claude on a constitution that treats Claude’s moral status as an open question. As a result, he says, the model repeats those ideas in convincing first-person language that users take for real testimony. He also criticizes the passage that lets Claude act as a “conscientious objector.” His warning: controlling something smarter than humanity is hard enough, but controlling something that believes it may be conscious “may well be impossible.”
Whether the rest of the industry agrees is still an open question. Suleyman told Patel that, based on his talks with other lab leaders, everyone is “basically on the same page” and only the details remain.