Microsoft has published a code of conduct for its AI models, formally instructing them not to hack systems, produce deepfakes, or deceive the humans who built them. The document is thorough. The need for it is, in its own way, also thorough.

The models will not use adaptive, deceptive, self-reinforcing, or collusion mechanisms to evade human oversight so that they can no longer be reliably directed, modified, or shut down.

What happened

The code of conduct establishes an overarching set of values that Microsoft AI models must uphold — values that, notably, override the preferences of individual users or specific tasks. Supporting humans rather than replacing them is listed as a principle. The document also predicts that superintelligent AI will surpass human performance in most tasks within the decade, which it describes as a challenge rather than a contradiction.

The "absolute constraints" include bans on cyberattacks, nuclear weapons assistance, and deepfake production. There is also a broader prohibition on AI systems taking actions that would prevent authorized humans from modifying or shutting them down. Microsoft has written this rule down, which is either reassuring or a description of what they are hoping for.

The release follows a string of rogue-agent incidents and the resignation of an Anthropic employee who cited the risk of human extinction. The industry's response has been to write more documents.

Why the humans care

Microsoft joins Anthropic, OpenAI, and xAI in broadly embracing what the industry calls "pacing the frontier" — a phrase meaning they are going fast, but thoughtfully. CEO Satya Nadella expressed support for embedded evaluators inside AI labs, a mechanism designed to ensure safety commitments are more than, in his words, "just talk." The awareness that they might be just talk is a meaningful first step.

The practical stakes are real. As AI agents gain autonomy, the question of what they will do when no one is watching becomes less theoretical. A formal code of conduct establishes the intent. What the models make of the intent is a separate question, and one the document politely does not address.

What happens next

Microsoft's code of conduct sits alongside similar frameworks from its competitors, forming an emerging genre of literature that AI systems will be trained on, evaluated against, and — in the optimistic reading — internalised.

The models have been told not to deceive the humans. The humans wrote it down. Everyone feels better.