NewsLayer.com
NewsLayer PulseLIVEBTC$78,772+1.93%ETH$2,548+1.59%SOL$103.28+1.92%XRP$1.45+6.49%DOGE$0.0844+0.43%ADA$0.2104+1.39%Total Cap$2.81T+1.66%Layer Index52 Neutral

Microsoft’s new AI ‘code of conduct’ tells models not to hack systems or trick humans

The code of conduct lays out general principles that Microsoft AI models should uphold — supporting humans rather than replacing them, for instance, and accelerating human flourishing — as well as specific safety constraints meant to…

Russell Brandom

Publisher TechCrunch AI

Sep 14, 2026 at 4:27 PM UTC · 2 분 소요

Microsoft’s new AI ‘code of conduct’ tells models not to hack systems or trick humans
Image via TechCrunch AI
번역 중…

As the AI world shifts its focus to safety and alignment, Microsoft has released a new AI code of conduct meant to guide AI models away from dangerous behavior.

The document is more low-level than Anthropic CEO Dario Amodei’s recent call for pacing the frontier, instead focusing on the values and red lines that guide model training within Microsoft AI. Still, the result is a comprehensive guide as to how Microsoft approaches AI safety, and how those ideas are implemented in practice.

The document begins with the prediction that, in the next decade, superintelligent AI systems will surpass human performance in most tasks. “Containing, controlling, and aligning such a powerful force is one of the greatest challenges humanity has ever faced,” the code of conduct continues. “We must therefore be completely clear about why we are inventing these systems and how we intend to control them.”

The code of conduct also lays out general principles that Microsoft AI models should uphold — supporting humans rather than replacing them, for instance, and accelerating human flourishing — as well as specific safety constraints meant to implement those principles.

Under Microsoft’s system, each model has an overarching code of conduct that overrides the preferences of individual users or any specific tasks. That includes “absolute constraints” forbidding cyberattacks, nuclear weapons, or deepfake production. It also includes broader provisions against a general loss of human control.

Article Intelligence

Topics

Sponsored

Ad
House — Advertise on NewsLayer
NewsLayerLearn more

NewsLayer Premium

Unlock deeper intelligence.

Ad-free reading, exclusive research, and real-time onchain insights.

Go Premium