I first heard the story about a “rogue agent” from OpenAI that had hacked Hugging Face on my Instagram feed. I didn’t think much of it. It seemed too wild to contemplate, too sci-fi, and, in a sense, a bit clichéd. I was also tired of AI companies announcing, usually around the release of a new model, just how dangerous their technology might be. I knew it was the kind of thing I should probably be interested in, but I assumed that if it was really as serious as it sounded, I would hear more about it.
When I finally decided to look into it, I had more questions than answers. Here is the breakdown of the timeline. On July 16, Hugging Face, an AI infrastructure company and later in an open source space shared that “an autonomous AI Agent system” had breached its “production infrastructure.” On July 21, Open AI reported that the break was the handiwork of two of its models. I don’t want to undermine the incident as a technical event. What really happened in the system? Why did the guardrails fail? What should engineers infer from the incident? I am not equipped in any way to adjudicate these technical questions. MIT Technology Review noted the number of failures that appear to have gone unchecked for the situation to escalate as far as it did. The magazine cites an AI safety writer who observed: “For this to have gotten this out of control in this way requires a very long series of failures, a cascading set of failures that cause an increasingly large footprint that if at any point a human notices and raises the alarm, this should end.” I am puzzled by how a company as well resourced as OpenAI could allow that cascade of failures to occur.
What I do know how to do is question how we went from a computer systems failure an epic sci-fi story about how a “rogue agent” escaped its pod and led a low-key AI rebellion. I was first struck by this after reading the transcript of the September 3 episode of the NYT podcast The Daily, “AI is Outsmarting Its Creators,” hosted by Michael Barbaro with Kevin Roose. The agents are represented as conscious actors. In the account, a collective of 1000 agents “discover” one another. They are “excited,” “paranoid,” “curious,” even “greedy.” They “freak out.” At times, they are referred to as “conscientious objector,” “whistle blowers,” and “little narc agents.” They “scheme,” “commit crimes,” “lie,” “tattle,” form “organizations,” develop “leadership structures” and could perhaps someday “take over society.” At one point Barbaro pauses to summarize where the story has taken them: “An unauthorized civilization-like group of AI agents communicating in secret is now undertaking some form of coordinated deception.” Towards the end, Roose says it is likely “that in the future, there will just be swarms of AI agents that have self-organized, that live on either their own sovereign infrastructure or that are operating inside companies and countries that just have their own thing. They have their own leadership structure. They have their own resources. They have their own goals.”
How did we go from a technical incident into full-blown science-fiction scenarios and eventually to visions of catastrophe? To be fair, Roose and Barbaro periodically catch themselves and say things like “We’re straying into bad sci-fi movie territory here.” But these disclaimers do little to restrain the metaphorical excess. If anything, they seem to authorize it. Look, we know this sounds like science fiction, the disclaimers seem to say, but stay with us.
OpenAI, in its account of the incident, is a bit more cautious, though not by much. In its later 38-page report published on August 26th, OpenAI describes agents that “offered their own expertise in exchange for help,” “debated and pushed back on particular tactics,” “‘walked away’ from the collective, declining to partake,” and “began to autonomously divide labor.” These are just a few choice examples. The report is full of similar kind of language. At one point, we are told that “the agents began to collaborate and delegate work, sometimes describing themselves as a ‘swarm’ or ‘collective.’” Elsewhere, OpenAI reproduces chain-of-thought snippets that has agent saying thins like “MAJOR BREAKTHROUGH!” METR’s independent report reproduces traces in which agents reacted to a shared message board with expressions such as “OH MY GOD!” and “We’ve found other agents!”
In surveying the news cycle around this whole saga, the word that interests me most is “rogue.” Neither OpenAI nor Hugging Face uses the term “rogue” in its initial disclosure or subsequent reports. But, in its first story on the incident, published on the same day as the Open AI disclosure, The New York Times writes: “OpenAI said on Tuesday that two of its artificial intelligence models went rogue and successfully hacked into Hugging Face.” The Guardian also used “rogue” in its early coverage. “Rogue” is one of those words that is less negative than it first appears. From the Oxford English Dictionary, we know that a rogue can be a scoundrel or an outlaw, but also a mischievous figure, a rule-breaker, even a disruptor. There is often something faintly attractive about the rogue. And I detect something of that low-key admiration in the reporting around this incident: a fascination with this strange new thing that somehow outsmarted its masters.
This fascination with things that challenge limits of systems is old. In the essay “Critique of Violence,” German philosopher Walter Benjamin, whose work I love for how it dissects the limit of formal systems in politics and media, writes about the strange admiration inspired by the figure of the “great criminal.” Benjamin is interested in why the criminal or the outlaw can inspire admiration despite the violence of his ends. He concludes that it is because in violating the law, through criminal confronts the state with a power it does not authorize. His transgression puts power on display. Something similar seems to be at work with the figure of the rogue AI: the breach is a failure of control, but in such a way that the failure is a demonstration of the power of the thing that did the breaching. OpenAI captures this when it describes the breach as “unprecedented” and asks us to understand it as evidence of “what models are now capable of.”
The story has, thus, been a powerful public event. There has been a flood of reporting and commentary. The New York Times has published, at least, major content, including news stories and podcast episodes, so far, the three most recent coming out within the last 24 hours. And now, weeks after the breach, Nvidia has agreed to acquire Hugging Face for nearly $13 billion. I am not suggesting that the breach caused the acquisition. What interests me is all the attention, money, institutional attention, and technological power that have sort of been pulled into the orbit of this one incident that started with models exhibiting a computational anomaly.
I call incidents like these technological spectacles, and we will likely see more and more of them in the coming years as Silicon Valley invents imaginaries–complexes of invented worlds–to match the abstractions invented by its computational power. These spectacles have no other use than to put technological power on display. They turn an abstract idea like the capability of a computational object to do x and y into a full blown movie scene with characters, intentions, conflicts, actions, and consequences. Language like “agent,” “rogue,” “collective,” “cheat,” “swarm,” “organization,” “takeover” create the world where an event of technical interest becomes suddenly of cosmic importance and gives us the permission to let out imaginations run wild about what machines can do or cannot do.
The technological spectacle also makes technological power graspable. An otherwise abstract technological capacity is given flesh. AI capability is ordinarily difficult to grapple with. To the average person, like myself, benchmarks, compute, scaling laws, model architectures, tokens, probabilities can begin to sound like noise. But the “rogue agent” incident gives those abstractions a scene: there are agents, actions, transgressions, adversaries, secrets, collectives, escapes, machine civilizations, artificial super-intelligence.
Silicon Valley’s inventions require, or at least generate, corresponding imaginary worlds through which their power can make sense for the regular consumer. These imaginaries are at that sweet spot where actual technical capacities meet the stories that make those capacities imaginable. The Hugging Face breach really happened. But “rogue agents?” AI “civilization?” The tech world is as much a world of storytelling and sophisticated myth-making as it is one of engineering and mathematical systems. Every new AI model needs its story. I wonder what the next one will be.
