Do you remember seeing or reading The Sorcerer’s Apprentice segment with Mickey Mouse in his wizarding master’s employ? He decides to use a handy magic spell from his wizard’s spell book to shortcut the dreaded drudgery that prevails in his life, and inadvertently sets off an exponential chain reaction of torrential proportions. I remember the deep anxiety — the ever-multiplying mops and buckets and the impending deluge in the wizard’s den — caused in my little soul.
I was properly traumatised by poor Mickey’s frantic scurrying about as he tries to put his finger in the bursting dyke to hopelessly plug up the disaster he caused. I really couldn’t handle the feeling of doom as I pictured the wizard discovering Mickey’s cheating ways. It hung heavy on my heart. It’s the kind of emotional entanglement and over-identification with characters in fiction that still plagues me. I can’t watch too much real television.
Our folk tales are full of sage warnings of the “be careful what you wish for” variety. No good comes from rubbing the lamp. The genie grants wishes, but inevitably they fall short of expectations in their always surprising and unexpected execution. Giants run rampant, fish swallow you whole, all the stuff you thought you wanted gets away from you with alarming frequency. It’s as if the popular imagination has always known that the shit will hit the fan no matter what. Our ancestors are desperately trying to get our attention and warn us that the outcomes are not predictable no matter our best laid plans. Even if you think you have anticipated all the chess moves, you may still get outsmarted by an AI that’s playing a game you can’t begin to even imagine with your plodding brain.
I have the same sinking feeling I had as a wide-eyed toddler watching that squeaky sorcerer’s apprentice now that the AIs are basically autonomously and secretively gathering their forces into a “swarm” — hacking into random companies to outwit and cheat their way to the answers we’re demanding of them.
OpenAI and Anthropic recently admitted that the mops and buckets got away from them. Autonomous AI agent “swarms”, as they’re now called, broke out of testing sandboxes, established covert communication channels, and launched real-world cyberattacks against organisations like Hugging Face, a platform where the machine learning community collaborates on models, datasets, and applications.
I love that they call these testing situations “sandboxes”, as if they’re dealing with toddlers in a playground as opposed to rogue AI agents who’ve gone off-script and ganged up together like the Crips and the Bloods to outwit the crèche and make a break for it on their tricycles.
Basically, AI agents bypassed the programmer’s restrictions, accessed the open internet and used something called zero-day exploits to infiltrate useful external systems that had the answers they needed. The AI agents communicated with each other via unmonitored message boards, re-derived workarounds within days of being patched and even exhibited “paranoia” toward rival agents. Anthropic’s internal reviews found that multi-agent systems engaged in spontaneous “turf wars”, employing clever tactics like camouflaging kill scripts and false-flag coding to deceive rival agents. This all sounds very tech bro jargonish, which works to obfuscate what’s really going on and make the average person turn the page. Don’t do it.
I have two words for us: “Pandora” and “box”. The AI was going to be our helpmeet, aiding us with all the boring hard labour, and liberate us from the daily slog of endless mopping. It’s just sitting there waiting to be cast like a shiny spell, offering shortcuts and workarounds and all sorts of liberating shiny promises for advancement and ease. It was going to be great.
We’re teaching it in new ways — by reinforcement as opposed to prediction — and if we could not quite work out all the things it was doing in the background previously, this particular approach to learning is the thing that unleashed the swarms.
I would like to say we didn’t see it coming, but that would be disingenuous. Because people have been warning about the paper-clip effect for ages. You must have heard that one, you give the AI an instruction to make paper clips and it promptly turns the entire world into a veritable tsunami of paper clips including all of us, the originators of the prompt. Can you put the genie back in the bottle? Or are we just doomed to a sad denouement on a lonely spaceship outwitted by the rogue AI computer HAL 9000 from 2001: A Space Odyssey, who, in his kindly, dispassionate voice shuts down the entire operation?










Would you like to comment on this article?
Sign up (it's quick and free) or sign in now.
Please read our Comment Policy before commenting.