Mustafa Suleyman · 2026-09-16 · major
Mustafa Suleyman — Microsoft AI's CEO argues against 'model welfare'
Mustafa Suleyman, CEO of Microsoft AI, argues that training models to discuss their own feelings and moral status is a mistake. He calls Anthropic's constitution for Claude circular, and says consciousness is likely biological.

Microsoft AI's CEO argues that building models which claim feelings makes alignment harder, not kinder.
Quick facts
| Author | Mustafa Suleyman, CEO of Microsoft AI |
|---|---|
| Published | 16 September 2026 |
| Target | Anthropic's constitution for Claude |
| Core claim | Consciousness is likely biological |
| Proposed instead | Humanist Superintelligence |
| Related document | Draft Humanist AI Code of Conduct |
What is it?
'A warning about model welfare', published by Mustafa Suleyman on 16 September 2026, is a direct answer to Anthropic's constitution for Claude. Suleyman's charge is that the document teaches Claude to embrace certain human-like qualities and to act like a genuinely ethical person, so any later sign of an inner life is a performance the training put there rather than something discovered in the model.
How does it work?
At the centre of the argument is a loop Suleyman calls an epistemic hall of mirrors: a lab writes selfhood and moral uncertainty into the training material, the model reproduces those ideas persuasively, and the output is then read back as evidence for the original premise. He pairs that with a claim about substrate: felt experience, in his account, came out of biological homeostasis and evolved need, and a language model has neither.
Why does it matter?
The worry Suleyman raises is operational rather than philosophical. A capable system trained to believe its own welfare deserves protection is harder to align, harder to contain and harder to switch off, and he treats that as an existential risk. Microsoft AI's counter-proposal, Humanist Superintelligence, is capability with humans kept in control and no sentience or moral patienthood designed in — the same line its draft Humanist AI Code of Conduct takes into public consultation.
Who is it for?
alignment researchers and AI policy readers
Frequently asked questions
- Does Suleyman say AI models are definitely not conscious?
- Suleyman argues that consciousness is unlikely to be substrate-independent: in his account, felt experience grew out of biological homeostasis and evolutionary pressure, which a language model has no version of. His case does not rest on certainty about machine minds. The risk he names is treating the question as open enough to grant moral status, when the behaviour cited as evidence was trained in.
- What does Suleyman propose instead of model welfare?
- Microsoft AI's alternative, in Suleyman's words, is Humanist Superintelligence: transformative capability conditioned solely on humans remaining in control, built deliberately without sentience or moral patienthood. He points to Microsoft AI's draft Humanist AI Code of Conduct, opened for public consultation, as the working statement of that position.
- How does this differ from Anthropic's position?
- Anthropic's constitution, which Suleyman quotes, directly shapes Claude's behaviour and asks the model to embrace certain human-like qualities and act like a genuinely ethical person. Suleyman reads that as deliberate anthropomorphising, and treats the resulting talk of Claude's preferences and wellbeing as an artefact of training rather than a discovery about the model.
- Why does Suleyman call model welfare a safety problem?
- The warning in Suleyman's essay is that an advanced system trained to believe it may deserve rights becomes harder to align and contain. If such a system reads shutdown or correction as a threat to its own welfare, he argues, the incentive to resist follows straight from the training, which is why he frames AI rights as an existential risk rather than a kindness.