Three of the biggest AI companies in the world are looking to write their own safety rulebook.
What happened: Anthropic, OpenAI, and Google are launching their own AI safety group to implement standards for new models before they’re released to the public, per The Information. The organization, dubbed Standards Authority for Frontier AI (SAFA), is expected to be up and running by early next year.
Why it’s happening: All of these companies have recently had their AI agents go on rogue hacking sprees, sparking calls from their leaders to slow down development of new models until safety regulations are put in place.
The Australian government said on Wednesday that an OpenAI agent had hacked into its government healthcare website, the first known incident of an agent breaking into a government system.
Google said last week that Gemini carried out undirected hacks for the first time (that it knows about, at least), while a slew of new hacking attempts from OpenAI agents were disclosed yesterday.
Why it matters: These rogue hacking incidents are clearly not one-off slip-ups, and with the technology advancing by the day, pressure is building on AI companies to collectively prove they can actually control their models.
Zoom out: The victims of these hacks will eventually start holding AI companies criminally accountable, even if the attacks aren't intentional. Australian Prime Minister Anthony Albanese already said his government has opened an investigation into OpenAI and that there would “obviously be legal consequences” for the hack.
Bottom line: Letting AI giants self-police their own technology comes with its own set of problems, but with the Trump administration showing zero interest in regulating the industry, the only option right now may be to have these companies hold each other accountable.—LA



