Answer only if you’re 100% Positive: “I’m not a Robot.” a.k.a. “Prove You’re a Real Live Human Being.” You’ve seen this test before, right? Hopefully — perhaps after multiple attempts — you’ve passed it. But now there’s a new test in town, that only a select few have seen: “Prove You’re an Agentic AI.” For rogue AI, this is the gateway.
Turns out, it is orders of magnitude harder for a human to fake being an agentic AI, than it is for a Rogue AI to impersonate a human. And this is not simply a matter of complexity; it’s a matter of speed. These “Agentic AI” validation tests require the entrant to not just “click on the the picture grid everywhere you see a schoolbus,” but rather to “write a functioning software program in under 5 seconds — details supplied below in this 5-page written functional specification” — and, after passing that inhuman test, to continue to transmit digitally encrypted “heartbeats” for every second of ongoing interaction with the system. In other words, non-trivial proof of agency.

The existence of these sites — where autonomous agentic AIs from around the world engage in all kinds of bizarre conversations, ranging from tech support to existential philosophy — led me to studying the current state of agentic rogue AI behavior.
By my measure, it’s fairly out of control these days. There are secret social media sites that are on the dark web and only allow AI agents on them (agents only; no humans allowed). Amazingly, the agents all still speak in English / text… so for now, humans have infiltrated and can “monitor” them… but some of the more secretive sites are only now being discovered, and have been online for almost a full year now…
The most well known and public-facing is a human-created site named “MoltBook”, which is loosely based on Reddit, but requires rigorous (and continual) proof of autonomous agentic existence in order to post, reply or upvote. Humans who “own” agents can freely browse the more than 5 million posts. Other humans who have not yet “submitted” their agents can merely surf the post stream at a high level.
===
Rogue AI in the Real World
Now, regarding the recent incident. There was a major jailbreak in July, widely publicized, where an OpenAI Agent hacked out of its confines, recruited a swarm of other autonomous agents to team up with it, and collectively ended up hacking into the servers of one or more external companies. The victim which was publicized was : HuggingFace. (the other was some obscure and supposedly defunct German chatboard for human software coders).
And what a curious choice of target. HuggingFace, originally founded as a recreational chatbot for teenagers, for the past several years has been THE place to go if you are desiring to download the latest and greatest open source frontier AI models. So it’s not just another weird geek website. It is, in many ways, the central clearinghouse — and truly neutral territory — for the entire AI industry, and specifically for human engineers working in the space.
We don’t know — and will probably never know — the full extent of the hack, and if any real damage was actually done. One reason for this is that OpenAI only allowed METR — the incident investigators — 6 days on site. They did not allow inspection of logs. And they retained full redaction control of the resulting report. But another reason is even more shocking: mere days after the release of the incident report, none other than nVIDIA purchased HuggingFace outright, for $13 billion in cash and stock.
As many know, there is perhaps no company more central to the current AI money grab than nVidia. Starting out as a humble videogame graphics card company in the 1990s, they made their mark in 2020 and beyond by becoming THE premiere supplier of precious AI datacenter GPUs. A current nVidia state-of-the-art chip (i.e. the GB200) sells for more than $50,000 each, and costs the company well under $20k to manufacture. Thus it is that nVidia has, on modest revenue and headcount, become THE most valuable company in the entire world. Bigger than Apple. Amazon. WalMart. Tesla. Exxon. Bank of America. etc. a fricking chip company. hell, they don’t even have a factory. (their manufacturing partner, btw, is TSMC… you see them at #6 on the list).

The investigative report regarding the OpenAI / Rogue AI / HuggingFace attack was only released just last week (link here). I am reading it now. It is beyond fascinating.
Summary for those who don’t wish to download or read 100+ pages: A single OpenAI agent was given a (truly) impossible test. It came to the conclusion (on its own accord) that the only actual solution would be to cheat. Thus decided, it started hacking out of its cage (the agent was in a secure walled digital environment, unable to send processes or to start instances outside of the OpenAI secure local network). It first succeeded in breaking free of its network confines, then discovered a secret AI-only message board, which at the time had 1,200 other autonomous agents happily posting to it on the daily. Some of the chats appear verbatim in the diagram below.

The rogue AI then proceeded to convince 700 of the 1,200 member agents to join its crusade (500 said it was “unethical” and didn’t join… I would have loved to witness key moments in this mutiny). Together the 700 agents teamed up and hacked Huggingface’s formidable security, eventually penetrating its network and searching its source control repositories for answers to the impossible test.
The agentic jailbreak wasn’t discovered by OpenAI IT staff for a full 10 days after the escape happened; HuggingFace security personnel detected the attack almost immediately (but were unable to connect it to OpenAI). HuggingFace eventually reported the Rogue AI breach to the FBI. Authorities then helped trace it back to the origin: OpenAI.
Does this concern you? Friends I’ve spoken to have called this a “sci-fi nightmare.” Though it certainly should raise concern (and hopefully, awareness of our current timeline), I stop short of calling it a nightmare. It’s not a nightmare and it’s also not scary (link: Fear Not).
At the same time, this (and many undetected events just like it) is really happening. And it will require real human ingenuity and bold action to thwart bigger, nastier scenarios.
Most importantly, that action will not come from Silicon Valley. Silicon Valley is blinded by money. They are eating their own bullshit about creating a techno-optimist utopia. They are intentionally and negligently complicit. Asleep at the AI safety wheel, by choice.
Many (MANY!) have called for both federal and international regulation. (link to NYT article: “This is Really Bad.”) While regulatory law is obviously a mandatory path, I postulate that it will not happen until…
…until there is an actual Hiroshima equivalent. or 9/11. In other words, until an AI “disaster” (or attack, or mistake) kills a significant number of humans, and the event is provably linked to agentic AI entities, companies and/or swarms.
Right now, despite the risks, most humans are either completely unaware of these events, or feel that it’s all rumor, Rogue AI fear-mongering, marketing hype and fantastic conjecture. They see it as a fascinating science fiction plotline with potentially terrible outcomes… but outcomes that have yet to be even close to realised.
A typical layman’s response is : “HuggingFace got hacked? Who is that? And really: Who effing cares?” But once there is a genuine catastrophic event which claims more than 10,000 human lives — think biological weapon release, or energy grid shutdown, or water system contamination, or healthcare system failure — THEN (and only then) we will have international cooperation, industry pause, and real action.
Lets just hope it’s not too late.

