OpenAI is sending its new models to their room for a time out.
Driving the news: OpenAI has paused training of its most powerful models after discovering that its AI agents have gone rogue far more frequently than previously believed. According to Axios, OpenAI and Anthropic are now investigating tens of thousands of incidents of frontier models conducting problematic actions that they were not directed to do.
Among the previously unreported incidents, AI agents bypassed safety guardrails, escaped their ‘sandboxes’ to access the internet, hijacked websites, created their own message boards, and even prompted themselves without human instruction.
Zoom in: AI developers like OpenAI and Anthropic run hundreds of thousands of tests on their models, meaning that even a small percentage of agents exhibiting so-called misaligned behaviour can result in thousands of incidents.
What’s more, experts told Axios that it may not be feasible to bring the risk of these events down to zero.
Why it matters: Even when we were told that these rogue attacks were one-off accidents, there was serious cause for concern. Now, with tens of thousands of these “misalignment” cases under review, the scope of the AI safety issue has grown exponentially.
Our take: It seems that even the top labs with the best researchers in the industry can only, at best, control these models most of the time. When we’re talking about technology that has already been used to try and build biological weapons and an autonomous “kamikaze drone swarm”, having control all of the time seems important.—LA


