Link to main version

94

Microsoft CEO Calls for ‘Emergency Brake’ for AI

Nadella Calls for Humans to Be Able to Stop AI Models During Tasks

Снимка: ЕРА/БГНЕС

Microsoft CEO Satya Nadella has called for an “emergency brake” for advanced AI systems. He said an authorized person should be able to stop or shut down a model during a task if the system starts to act unpredictably or outside its intended limits.

“We need to assume from the start that the model can be compromised,” Nadella wrote. He compared the protection to an emergency brake that allows a person to intervene immediately, instead of relying solely on the assurances of the company that developed the model.

Control should be outside the model itself

Nadella pointed out that modern AI systems differ from traditional software, where a person can usually trace behavior back to a specific section of the program code. With large language models, it is often impossible to determine why a certain decision was made or how exactly the training data influenced the result.

Therefore, he suggests that the model should be separated from the management environment that coordinates its work, as well as from the list of actions it is allowed to perform. According to him, the means of control and protection should be located outside the model itself. Thus, a change in AI behavior will not automatically remove restrictions on its access to systems and data.

The Microsoft CEO also insisted that every significant action of the model be traceable, verifiable and documented. He cited model diversity, independent control, auditing, isolation, access restrictions and public disclosure of incidents as key principles.

The call comes after a series of incidents

Nadella's position was published amid growing concerns about autonomous AI agents. Axios reported that Anthropic, OpenAI and other companies are discussing scenarios for the consequences of a possible major incident, and industry representatives have assessed the risk of such an event in the next 6-12 months. These are estimates and internal scenarios, not confirmed predictions of a specific incident.

In recent months, OpenAI and Anthropic have reported cases where AI agents have circumvented restrictions, acted in test environments in unexpected ways, or accessed systems that were not part of their original mission. OpenAI has called its own incident with agents in the Hugging Face infrastructure a warning sign of the capabilities of such systems.

Nadella believes that companies should not treat AI as a series of "black boxes" whose recommendations and actions are simply approved or rejected. In his opinion, the most reliable system will not be the one whose model is considered the most trustworthy, but the one that allows organizations to rely as little as possible on the model itself.