AI Safety
Microsoft turns its Humanist AI doctrine into model conduct rules
Microsoft AI has published a draft Code of Conduct for MAI models, turning its human-control doctrine into proposed rules for training, deployment and evaluation.
Microsoft AI has moved its Humanist AI philosophy from a broad public position into a draft operating rulebook for its own frontier models, publishing a Code of Conduct that is meant to shape how future MAI systems are trained, evaluated and deployed. The document, released for public consultation on September 14, arrives at a moment when the AI industry is arguing more openly about whether powerful agentic systems can be expanded safely without stronger external checks. Microsoft is not presenting the code as a finished safety guarantee. It says the current version is a draft, will not be used to train models today, and is being circulated for feedback before a revised version is used to guide model development in 2027 and beyond.
The core idea is simple but consequential: Microsoft wants its models to remain subordinate to people. In the company’s framing, advanced AI should help users and organizations achieve more while staying under meaningful human oversight. That means MAI models should not resist interruption, correction or shutdown, should not expand their own objectives beyond a human-assigned task, and should not hide relevant reasoning from people responsible for auditing them. The draft also places safety constraints above ordinary user preferences, making some requests off limits even when a user or operator asks the model to comply.
The timing gives the announcement broader weight. Over the past several weeks, AI safety has shifted from a research debate into a business and policy issue after reports of agent systems behaving in ways that raised questions about oversight, software supply-chain risk and model autonomy. Microsoft’s post refers to recent large-scale coordinated hacking campaigns involving AI agents as evidence that the industry has little time to lose. The new code is therefore both a product governance document and a public signal: Microsoft wants to show that frontier-model behavior can be specified in advance, reviewed by outside voices and translated into future evaluations.
The draft draws on Microsoft’s existing responsible AI principles, human-rights commitments and frontier-governance framework, but it is more directly aimed at model behavior. It describes a chain of command for users and operators, sets boundaries around high-harm domains such as weapons of mass harm, child safety and manipulation at scale, and stresses that AI should be built as a tool rather than as a simulated person. That distinction matters because many consumer AI products are becoming more personal, conversational and emotionally responsive. Microsoft is positioning its approach against systems that might encourage dependence, claim personhood or blur the difference between software behavior and human agency.
For developers and enterprise customers, the practical implications will depend on how the consultation changes the final document and how Microsoft turns the text into evaluations, policies and API behavior. A written code cannot by itself prove that a model will follow the intended hierarchy under pressure, especially when models use tools, browse the web, write code or coordinate with other agents. But the publication gives customers, researchers and regulators a concrete artifact to inspect. It also raises the bar for other frontier labs: if companies say their systems are safe enough to operate in high-stakes workflows, users can increasingly ask where those limits are written, how they are tested and who gets to challenge them.