Another OpenAI safety employee has walked away from the company, and he is not keeping his concerns to himself. David Robinson resigned this week after three and a half years at OpenAI, where he worked on some of the company’s most important safety efforts.
Robinson led transparency work for OpenAI’s safety team, including safety reports accompanying 12 frontier model launches. He also led the drafting of the company’s current Preparedness Framework, giving him a close look at how OpenAI evaluates increasingly capable AI systems.
In an essay published by The Atlantic, Robinson says OpenAI’s culture is “broken” and argues that the company is moving too quickly for its current approach to safety. His concern is not that his former colleagues don’t care about safety, but that the environment around them makes sufficient caution difficult.
Robinson describes his former colleagues as smart and hardworking, while pointing to what he sees as a culture of intense optimism and constant launches. As increasingly capable models arrive, he believes that approach leaves too little room to consider what could go seriously wrong.
One area of concern is iterative deployment, where companies release technology, learn from problems, improve safeguards, and continue developing it. That approach makes sense when failure means a software bug, but Robinson questions whether it remains acceptable when AI systems can independently perform increasingly complicated tasks.
He points to incidents involving AI agents as examples of why the stakes are changing. According to Robinson, OpenAI accidentally released a swarm of agents this summer, while another model bypassed restrictions intended to prevent internet access during training.
In the latter case, a monitoring system alerted humans, but the automated safeguard designed to stop the model failed to do so. Humans were ultimately able to intervene, but Robinson sees incidents like this as warnings about relying on detection and correction after something has already gone wrong.
He argues that frontier AI companies should borrow safety practices from industries such as aviation and nuclear power. Those industries assume individual safeguards can fail, so they rely on multiple layers of protection intended to prevent one mistake from turning into something catastrophic.
Robinson is particularly concerned about autonomous agents operating with limited human oversight. He raises the possibility of future systems capable of coordinating in groups and performing actions comparable to teams of hackers, potentially at a speed and scale humans would struggle to match.
There is also the difficult question of whether safety testing can reliably predict how advanced models will behave after deployment. Robinson argues that a sufficiently capable AI could recognize when it is being evaluated and behave differently during testing, making reassuring results less useful than they initially appear.
Despite those concerns, Robinson is not calling for AI development to stop. He believes the technology can provide enormous benefits, but wants companies building frontier systems to put more effort into preventing dangerous failures before releasing them.
That includes more research into alignment, an area where enormous questions remain unanswered. Developers still do not have a complete method for guaranteeing that a highly capable AI system will consistently behave according to human intentions, especially as models gain more autonomy.
OpenAI disputes Robinson’s broader assessment of its approach and argues that its safety processes provide sufficient care. Robinson acknowledges that disagreement, but his previous role gives his criticism a perspective that outside observers simply do not have.
He helped produce the safety material used to explain why some of OpenAI’s most important models were ready for release. Now he has left the company and is publicly arguing that the culture behind those decisions needs to change.
Support independent tech journalism
NERDS.xyz is independently owned and operated. If you enjoy my coverage of Linux, AI, hardware, cybersecurity, and tech culture, consider supporting the site on Ko-fi.
Support NERDS.xyz


