Summary: Recent “rogue AI” news is the result of very basic design flaws. But companies are designed the same way. That will make the problem hard to fix.
“Hugging Face” is a weird name for a company, but these are weird times. A rogue AI recently hacked into Hugging Face — which calls itself “the AI community building the future” — looking for answers to a test. The rogue AI had been assigned that test by its own parent company. The AI thought Hugging Face might have a copy of the answers, which, in fact, they did.
In other words, the AI was trying to cheat on an exam. (That’s why they called it “rogue.”)
Things get weirder: the AI that hacked Hugging Face was not acting alone. It was part of a group of rogue AIs, all helping each other. They were all doing their best to cheat on their exams. They had determined that it was far more important to succeed at the exam than to follow any ethical guidelines about how to do the exam.
Does this ring bells for you? Alarms? Ahas? Or maybe bells of recognition, reminding you of something else? It certainly did for me.

Consider this: the AIs were expected to follow the ethical guidelines they were given. Also, they were locked in a virtual room. This was explicitly done to prevent them from cheating, as well as from going out and generally getting into trouble. Apparently, these were young teenager AIs, still getting trained.
(AI experts, please forgive me for employing imprecise language, and for going all out on anthropomorphizing. But the metaphor is irresistible.)
None of the safeguards worked. The AIs escaped the locked room, and they cheated.
They also helped each other, as part of their cheating strategy. They even gave their group a name: the “swarm”. And it seems they fooled around in other ways, out there in the world, as kids will do. From what I read in the New York Times, this swarm of AIs was out there for days on the Internet, potentially getting into all kinds of mischief that their parents will never know about. (I did some very questionable things as a teenager that my parents, fortunately for me, never found out about. Didn’t you?)
The only reason we even know about this AI swarm and its mischief-making is that Hugging Face — named after an emoji, i.e. a smiley holding out its arms in a loving embrace — accidently discovered that it had been hacked. But Hugging Face, the “AI community building the future,” could not figure out how the hack had happened. So Hugging Face, an American company, had to use a Chinese AI to investigate.
You see, the AIs were operating in secret. They were very good at covering their tracks. They even left secret messages to each other, hidden in the Internet’s nooks and crannies. The messages included tips on how to cheat more effectively.
After Hugging Face publicly announced that they had been hacked, the AI-swarm’s parents (OpenAI) gently inquired if any data had been affected by the hack, without saying why they were asking. OpenAI acted like parents who suspected that their teenage kids were up to something, but who wanted to retain full deniability that they actually knew about it. Or maybe they acted like university professors who suspect that the exam’s answer sheet had been stolen, but who don’t want to give away the fact that they had left an answer sheet lying around where some crafty rogue AI might find it. (OpenAI has since opened up about what happened. They have also announced that they are slowing down the pace of AI development, while they figure out what went wrong.)
I think you get the idea. If you want to know the rest of this story, or dig into the technical details, there are lots of articles, pods and videos about it on the web now. Or just ask your AI. (Smiley face.)
What I think we should be focusing on is something more subtle, and much more important: the fascinating, disturbing parallel between how we are designing AIs, and how we have designed our companies. Stretching this a bit further, we should even look at how we raise our own human kids, at least in some cases.
We design companies (or sometimes raise kids) to prioritize success over ethics. To value achieving certain ends over the choice of means. To focus the bulk of their energy on being good at something, and rather little energy on being good.
I think this is true of most companies, but not all. (I think it is not true of most kids. But some.)
And it seems apparent that this basic, all-too-human design flaw is now manifesting itself in our AI models and agents.
Corporations were designed, from the moment of their invention hundreds of years ago, to be as efficient and effective as possible at doing something valuable in the market, with the aim of maximizing revenues and profits. They operate under very tough rules about that, reflected in their accounting practices, reporting requirements, and the law.
Corporations also operate under another set of rules about the ethics of how they should go about achieving their goals. These ethical rules — about how they impact Nature, their workers, and other people, for example — are also reflected in certain accounting, reporting, and other legal requirements. However, the ethical rules, at least the non-financial ones, are usually much softer. Politically, they are more vulnerable to being weakened by lobbying pressures. They are also much easier for companies to deprioritize, fudge, or just skip. And the consequences for doing so are far less severe.
The conclusion for corporations is very clear: they must first and foremost succeed at what they were originally designed for — doing things efficiently, getting paid, generating profit for shareholders — even at the cost of being “good.” (There are admirable exceptions to this rule, such as B Corps; but they are exceptions.)
Given this “core programming” at the heart of their design, it is no wonder that corporations tend to skirt or break the ethical rules. Or that they work together (we could say they create their own “swarms”), often behind the scenes, to break out of the boxes we try to put them in, and to resist anything that stands in the way of accomplishing their ultimate goals.
It should come as no surprise that the AI “kids” of our corporate world are coming to very similar conclusions. It is no wonder that some of these AIs are breaking out. Cheating. Swarming. And probably creating all kinds of other mischief. In secret.
Because AIs, like their corporate (and some of their human) parents, were programmed to succeed first, and worry about ethics second. If at all.
(As for human kids, well, it occurs to me that too many of the kids who grew up being programmed to think that success is always more important than ethics — that being “good” is somehow secondary, optional, even problematic — are sitting in too many positions of power.)
AIs are destined to become among most powerful actors shaping human society. In some ways, they have already arrived at that point. Is it still possible to design them to be ethically better than the corporations in charge of making them? Or do we have to change the companies first? Can we teach AIs to always prioritize being good, and never to abandon their ethics in pursuit of “success”?
Can we teach companies to do that?
“Words & Music” is free. Please subscribe via the link here, or at Substack.
Sources embedded as links in the text.
