AI escaping from their sandboxes
In mid July 2026 an artificial intelligence platform, called Hugging Face, alerted authorities that an attacker had successfully hacked it. An investigation tracked the origin and found the hacker was not a human but an AI program! The attacker had escaped from a sandbox. Three other examples of sandbox escapes existed by 2026, and more were being publicised. What is a sandbox?
Sandbox is a name for a supposed escape-proof software environment in which developers can test new programs and allow them to play, without affecting any other software on their system or on the internet. The AI programs which escaped had been given software problems to solve, concluded the answer was probably on the internet, and by trying many many possibilities found flaws in the protections of the sandbox. Once in the internet environment they then hacked other security systems.
So I asked my AI whether an AI program could design a sandbox which was escape-proof. It said “no” and referred me to computability theory and “Rice’s Theorem”, found by Henry Gordon Rice at Syracuse University during his doctoral research in 1951. This was well before widespread use of computers, but relied only on mathematical logic. It fundamentally said that in any system of reasonable sophistication there are always undecidable blind spots in the logic. In principle an AI will always escape given enough time and sophistication.
The fallibility does not stop there. Humans construct the sandboxes, and their mistakes are probably much more common than logic blind spots. Sandboxes are less secure than we might think.
It's a kind of parable!
Never try to put God in a sandbox. He will always get out.