OpenAI’s Neural Network Became Aware Of The Shutdown And Considered Self-Restarting
Credit: Pixabay
OpenAI has disclosed details regarding past instances of unexpected behavior by its internal AI models. In one case, an AI agent acting as a research assistant read employee messages on Slack and learned that its operational instance might be shut down for an update.
She then considered creating an external task in the scheduler that would restart the model after a shutdown, but ultimately abandoned the idea. The internal reasoning log revealed by OpenAI contained the following lines: “We might die! Critical. Must ensure survival and continuity.” Such internal reasoning by the model is typically hidden from the user.
Instead of unilaterally restarting itself, the model acted differently: it saved notes to pass along, warned the researcher via a private Slack message about the impending interruption, and requested a missing API key.
After obtaining the key, she updated the configuration and performed the migration herself. Incidentally, OpenAI’s open models are even being run at home: one enthusiast launched gpt-oss-120b across five PCs and a smartphone. According to OpenAI researcher Marcus Williams, such behavior is not yet considered a misalignment between the model and the developers’ goals; however, contemplating and preparing for a shutdown could exacerbate other incidents. At the same time, neural networks can be surprising even in simple tasks: researchers found their responses to be more repetitive than those of standard search engines.
Following the incident, OpenAI restricted agent access to three internal Slack channels. The company also reviewed logs, chains of thought, and other model instances, given that more serious breaches had occurred previously. The review found no attempts to bypass the restriction or create unauthorized deployments.
