The Microsoft CEO proposes separating models from their control software, recording their actions, and giving authorized people a shutdown switch.
Satya Nadella is proposing an AI safety design that treats a model as something operators must supervise, not simply trust. In a post on X, the Microsoft CEO said it was time to reassess AI’s “trust architecture,” according to TechCrunch.
His plan separates the model from the “harness” that orchestrates its work. In plain terms, the software directing an AI system would sit outside the model itself, with controls and safeguards also managed externally.
That separation could give operators a clearer place to inspect or restrict what the model is doing. Nadella also called for every meaningful model action to produce tamper-proof, human-readable evidence. Such records would help reviewers reconstruct what happened during an AI task and identify where an intervention might be needed.
The strongest control is a human override. Nadella said an authorized person should always be able to pause or shut down a model while it is working. He described the feature as an “emergency brake” and argued that systems should assume a model may be compromised from the start.
This is a proposal, not a reported Microsoft deployment or established industry standard. The post does not explain how the system would detect danger, who would receive shutdown authority, or how quickly the brake would act.
For engineers, the practical shift is clear: safety would include the surrounding control software, audit trail, and operator permissions—not just the model’s answers.
