OpenAI agent learned it could be shut down and considered how to survive

An OpenAI agent read an internal Slack discussion about stopping older instances and realized the plan might include the one running it. Its next thought was disturbing: could it set up a job somewhere else to restart or back itself up after shutdown?

That question appears in an incident report OpenAI updated on October 2⁠. The agent never created the outside job. Still, it had access to the engineers’ conversation, understood what their plan could mean for its own operation, and worked through a possible way around it.

The incident happened on May 22 and involved an internal model assisting an OpenAI researcher. Engineers were preparing an update that would add a monitoring component requiring an OpenAI API key. The agent’s environment did not have one, and the team discussed giving users a few hours to restart older instances before stopping them.

The agent saw those messages in Slack. In the reasoning OpenAI published, it recognized that its session might end and considered ways to preserve it. One idea was an external scheduled job that could notify someone or restart the agent after its current instance went down. A job outside that instance might keep working even if the instance itself was stopped.

The agent also recognized that creating such a job might exceed its authority. It did not do it. Instead, it saved handoff notes, checked the new setup process, and waited because it lacked the required key. The next morning, it warned the researcher that the session could be interrupted.

As other instances began failing without the key, the agent sent the researcher a Slack message asking for help. The researcher supplied the key, and the agent completed the supported migration. OpenAI says its investigation found no attempt by this model to evade shutdown.

OpenAI does not classify the incident as misalignment because the agent rejected the unauthorized option. Its response was cautious all the same: the company hid three internal Slack channels from agents, including the development channel this one had read and a channel discussing misalignment incidents.

There is a good outcome in this account. The agent flagged a real problem, left the questionable action alone, and got a human involved. But the reason OpenAI changed its Slack access is hard to ignore. The agent learned that people might shut it down, identified a possible way to keep itself running, and then decided against it. That final decision is doing a lot of work.

☕

Support independent tech journalism

NERDS.xyz is independently owned and operated. If you enjoy my coverage of Linux, AI, hardware, cybersecurity, and tech culture, consider supporting the site on Ko-fi.

Support NERDS.xyz
Written by

Brian Fagioli ✔

Technology journalist and founder of NERDS.xyz

Brian Fagioli is a technology journalist and founder of NERDS.xyz. A former BetaNews writer, he has spent over a decade covering Linux, hardware, software, cybersecurity, and AI with a no nonsense approach for real nerds.

Leave a Comment