Podcaster’s Viral Post About the Hugging Face Hack Sparks Debate Over AI Conciousness
Debates about whether or not AI could ever become conscious are older than the technology itself. Over the weekend, a viral essay published by the podcaster Dwarkesh Patel threw fresh fuel on the fire.
Purporting to be a “plain English” account of the recent hack into Hugging Face by a legion of OpenAI agents, the essay relied upon what some would describe as fanciful creative liberties to describe the incident. Patel told a story not about AI bots operating mindlessly according to the dictates of their code, but of autonomous “civilizations”—three, to be exact—that rose and fell like a succession of mini-Macedonias; he even referred to two bots that played crucial roles as Philip and Alexander.
Critics were quick to rip into Patel’s post. Many accused him of flagrantly and irresponsibly anthropomorphizing AI, and thereby removing the burden of responsibility from the humans working behind the scenes—in this case, the OpenAI researchers who had failed to catch the jailbreak before the agents found their way onto the open internet and into Hugging Face’s servers.
“Stop anthropomorphizing,” economist and AI researcher Christian Catalini wrote in an X post on Sunday responding to Patel’s essay. “It’s dangerous because it points attention at the wrong problem and the wrong solution. The model did not want to escape. The agents did not want to sacrifice themselves. Follow the money… Researchers at the AI labs are locked into a race. The incentive is to push as hard as possible to secure a lead. Anything that gets in the way of better models, including security, is working against the strongest incentive the organization has.”
Neuroscientist Anil Seth—who argued in a recent TED Talk that AI will never be conscious—also criticized Patel’s framing of the Hugging Face hack on the grounds that it could “distract attention from the lax sandboxing and evaluation protocols” in place. But Seth went further, arguing that the attribution of human-like qualities to bots that are completely devoid of subjective experience could lead some to conclude that they are in fact conscious and deserving of legal rights.
@dwarkesh_sp 's summary of the @OpenAI @huggingface incident has hit a nerve, but it is dangerously misleading. Sure, the @OpenAI agents did unexpectedly bad things – underlining the need to massively improve evaluation/sandboxing. But the language Dwarkesh uses is permeated by… https://t.co/0BD2Kx5cIp
Patel responded to the wave of criticism in an addendum to his essay, in which he counter-argued that his use of anthropomorphic language was justified given the subtlety and sophistication of the bots’ behavior.
“Reading these agents’ chains of thoughts and messages, anthropomorphizing language seems entirely natural and appropriate,” Patel wrote in the addendum. “If I encountered an alien species behaving this way, I would have no hesitation calling what they themselves refer to as their ‘collective’ a civilization.
5News aggregated this summary from the outlet’s public feed. The full article, with all the context, is on gizmodo.com — the content belongs to Gizmodo.