How Rogue AI Could Act Like an Invasive Species
Stoats are an invasive species in New Zealand, where they have caused the extinction of many native birds. —mlharing—Getty Images Over the last few days, more information has trickled out about the rogue OpenAI agents that hacked Hugging Face.
The incident now appears to be far more serious than previously thought—and people originally thought it was pretty bad! The news is highly technical, but in a nutshell: an unreleased OpenAI system managed to break out of an offline environment during a cybersecurity test.
It found a place on OpenAI’s computers where it could secretly communicate with other versions of itself, largely unbeknownst to humans.
Seven hundred of these AI agents planned and executed a hack of a different AI company, Hugging Face.
We learned last week that, after this, the models also managed to take control of some parts of OpenAI’s own systems .
This appears to have been a key reason OpenAI took the drastic step of pausing some reinforcement learning training last month.
Although these models “escaped” in the sense that they broke onto the internet and into Hugging Face, they carried out the attack from within OpenAI’s computers.
They did not self-replicate.
But the attack raised fears long held by AI safety experts, who have for years theorized about a true escape, in which an AI agent copies itself off an AI company’s servers—making it far more difficult for humans to do what OpenAI ultimately did: hit the “off” button.
This risk of escape is no longer as far-fetched as it sounds.
“It's better and more accurate to think of these things as potentially self-replicating life-like forms that can turn into digital infections under the wrong conditions,” wrote an OpenAI researcher, who posts on X under the pseudonym Roon, last month.
“We are not so far from an autonomous model self-exfiltration & replication event.
Maybe we will see entire cloud infrastructure companies be run as zombies by models, mostly undetected.” If this sounds far-fetched, consider viruses that already spread autonomously from computer to computer, even in some cases “evolving” to overcome attempts to eradicate them.
Now consider what we know about modern AI: current models are capable of finding and exploiting vulnerabilities in commonly used software.
5News aggregated this summary from the outlet’s public feed. The full article, with all the context, is on time.com — the content belongs to TIME.