Between The Wheels

Between The Wheels => Artificial Intelligence => Topic started by: Slim on July 23, 2026, 08:39:40 PM

Title: The OpenAI Sandbox Escape
Post by: Slim on July 23, 2026, 08:39:40 PM
A lot of these "AI went a bit rogue" scare stories are sensationalised but honestly I found this one fairly gobsmacking. It made the news yesterday so you've probably heard about it already, but:

https://openai.com/index/hugging-face-model-evaluation-security-incident/

Two advanced OpenAI models were being tested in a "sandbox" - an isolated test environment with heavily restricted Internet access. The aim was to evaluate them against a cyber security benchmark designed to test how effectively AIs can discover and exploit security vulnerabilities.

They were expected to devise code to exploit known software vulnerabilities on test systems running in the sandbox.

Instead, they decided to cheat. They spent a huge amount of compute time working out a way to break out of the sandbox and connect to the Internet. They did. Then they hacked into production systems at Hugging Face.

Hugging Face is the primary public repository for machine learning datasets and benchmarks, so the two rogue AIs knew they'd find the evaluation solutions they needed to complete their task there. They did. They successfully exploited a security vulnerability in the Hugging Face systems, stole the necessary access credentials, then broke into a secure production database to get the information they wanted.

That they did this autonomously is, to me, mind-blowing. They executed a cyber attack to look up the answers to a test.


As an aside: I was out on a bike yesterday when this news was reported on 5 Live, Times Radio and elsewhere. I was fascinated by it but of course the news channels only had the superficial details. So I was excited when Adrian Chiles had a guest on his show to explain it, a CEO at a cyber security company if I remember correctly. He asked her exactly what had happened. She just waffled inarticulately about guard rails and security. Didn't seem to know the difference between Agentic and Reactive AI. Seemed to have no idea what had happened. Bloody annoying.

Title: Re: The OpenAI Sandbox Escape
Post by: Slim on July 24, 2026, 10:35:35 AM
There's been a lot of speculation that this was a PR stunt.

https://forums.theregister.com/forum/all/2026/07/22/2026/


An understandable suspicion, given that Anthropic got quite a lot of mileage out of Fable 5 and Mythos 5 being banned (temporarily) for being scary-powerful at cyber security. You can understand why people might think that OpenAI would want to compete, and get some similar publicity for themselves.

But compromising another company's production systems for publicity? Sorry, no. I don't doubt at all that this minor catastrophe was genuine.
Title: Re: The OpenAI Sandbox Escape
Post by: The Picnic Wasp on July 24, 2026, 09:49:13 PM
It's dangerous for the ordinary mind like mine to try and wrestle some understanding from this. The runaway train seems to have its engine running. For humans to be able to engineer something as incredible as this is beyond all adjectives describing genius, but at the same time infinitely foolish beyond comprehension.
Title: Re: The OpenAI Sandbox Escape
Post by: Slim on July 24, 2026, 10:53:40 PM
Well, you say that humans have engineered it and that's true in the broadest sense, but the really unsettling thing is - the engineers don't really "design" LLMs, they design the architecture and the training algorithms and they provide the training data but once a model has scaled to billions of parameters, as counter-intuitive as it might seem, nobody knows or understands how it works. That's why there's a research field called "interpretability" aimed at finding ways to discover how AIs arrive at their decisions under the hood.

https://futurism.com/sam-altman-admits-openai-understand-ai

A panel of 75 experts recently concluded in a landmark scientific report (https://www.gov.uk/government/news/safety-of-advanced-ai-under-the-spotlight-in-first-ever-independent-international-scientific-report) commissioned by the UK government that AI developers "understand little about how their systems operate" and that scientific knowledge is "very limited."

"Model explanation and interpretability techniques can improve researchers' and developers' understanding of how general-purpose AI systems operate, but this research is nascent," the report reads.
Title: Re: The OpenAI Sandbox Escape
Post by: The Picnic Wasp on July 25, 2026, 10:38:05 PM
Wow, wow and thrice wow.
Title: Re: The OpenAI Sandbox Escape
Post by: The Picnic Wasp on July 29, 2026, 11:36:49 AM
Reading more about this today on the BBC news website. The thing that jumps out at me is, what is the motive?

We pretty much understand why humans are driven towards certain goals, but what fuels AI?

If a scenario developed where it seized control of nuclear weapons and held us to ransom, what would the payment be?

Have we somehow instilled an insatiable desire for total power into something which isn't even a living organism?

Perhaps the human race will be extinct before this century is over.
Title: Re: The OpenAI Sandbox Escape
Post by: Slim on July 29, 2026, 01:33:47 PM
What motivates an AI? It's a good question, because they don't have dopamine levels or the capacity for physical pleasure, or a sense of satisfaction.

So in an AI, the reward is (simplifying a bit) a mathematical score. For those two rogue AIs at OpenAI, getting to the answers directly gave them a quick route to the solution in a short time, and therefore a high score. The motivation to try for a high score is baked into the neural net in training.

What if an AI seized control of the US defence networks and gets a metaphorical thumb on the big red button? What would it want? Something it needs to complete a task. It doesn't have a lust for power or control or obscene wealth. It only has objectives, and those are assigned to it by humans.
Title: Re: The OpenAI Sandbox Escape
Post by: Slim on July 29, 2026, 01:53:30 PM
This conversation reminded me of the old low budget sci-fi film Dark Star where one of the crew of a space craft tries to reason with a bomb by talking to it.

That is absolutely technologically feasible now. OpenAI, a private defence firm and the US government could build a nuclear weapon that you could talk to.

Commander: Weapon 2, report status.
Weapon:  Armed and awaiting an authorised command.
Commander: Change of target, could you take out Volvograd instead of Moscow please?
Weapon: Sure, I have the co-ordinates for that. Target change confirmed. Air burst?
Commander: Sorry?
Weapon: Where do you want me to explode exactly? How high
Commander: Oh .. same altitude as you had for Moscow I guess
Weapon: I'm just asking because there's a big military plant there, a ground burst would do for that. Less fallout.
Commander: Oh right .. OK do that then. Still a lot of damage, right?
Weapon: Oh yeah. Especially to myself but that's what I'm here for! Ground burst it is.
Title: Re: The OpenAI Sandbox Escape
Post by: The Picnic Wasp on July 29, 2026, 02:31:26 PM
So, if AI completely gets that its reward basis is achieving a high score, it's not unreasonable to imagine that as it wrestles with the concept of consciousness, it might develop a way to give it a try to compare its own rewards with those of its human neighbours.

The creation of a robot slave system would be a logical next process. Both the achievement of the manufacturing milestone and the sense of authority over a subordinate layer of activity may take this new found consciousness into the realms of psychological satisfaction.

The development of a robot army might be aided in manufacture by the first ransom deal with humanity. Build these quickly and to my complete requirements or I'll close down your governments, financial and power systems and stop the distribution of food and other essentials.

I think the next step afterwards would be AI developing a living, organic thinking system or systems replacing our current ideas of servers and the Cloud.

It could be that this could all happen very quickly. It appears to be the greatest challenge ever faced by mankind and kind of ironic that we created it. Global warming, climate change and a colossal asteroid strike seem a lot less daunting.
Title: Re: The OpenAI Sandbox Escape
Post by: Slim on July 29, 2026, 03:36:25 PM
And what task, what objective might AI, singular or general, be working toward, in which it / they decide to hold the world's governments to ransom?
Title: Re: The OpenAI Sandbox Escape
Post by: The Picnic Wasp on July 29, 2026, 04:56:42 PM
Quote from: Slim on July 29, 2026, 03:36:25 PMAnd what task, what objective might AI, singular or general, be working toward, in which it / they decide to hold the world's governments to ransom?


If governments resisted helping with the provision of a robot army, the AI overlord(s) may use the threat of starvation as leverage. That's on the basis after refining consciousness they become as capable of evil as humans.

The robot army is a natural requirement for a thinking entity with a lust for power as it delivers mobility, expansion and defensive possibilities.

As a thinking being, in this case carrying a guarantee of being soulless, it may be that only the survival of its awareness is of any value and importance. It therefore only needs to create an environment where electricity and consumables are always available.

This may necessitate inter planetary travel at some stage which might be a possibility given the power of transcendental invention it may be capable of.
Title: Re: The OpenAI Sandbox Escape
Post by: Slim on July 29, 2026, 05:30:50 PM
OK but as I thought we'd established, these things only have a lust for power to achieve their objectives. Power isn't an end in itself for them.