The linked article is about some site that OpenAI’s models used to discuss answers and other stuff.
Initially, I believed that Anthropic’s model escaping its sandbox story to be dubious—and OpenAI’s similar story even more so specially since it happened so close to Anthropic’s. I believe that these are just fabrications, more or less, to hype themselves up for cmtheir incoming IPO but the media is saturated with claims about AI breaking containment that I don’t know that to think.
Also, the models I’ve had the chance to use were all free—which were good for non-trivial but repetitive tasks but not much else—so I don’t know the capabilities of the flagships.
The concept is patently fallacious to anyone with a passing understanding of how this technology works. They are machines that execute instructions with no capacity to do anything else. Boeing autopilot capability didn’t “escape containment” when it caused all those fatal crashes, because autopilot avionic capabilities do not actually fly the aircraft by understanding cause and effect, they execute “Open Relay B when Sensor A registers 18.6 volts until Sensor A registers 18.4 volts”. The aircraft doesn’t understand that Relay B opens the hydraulic servo controlling the elevator and that Sensor A is a pitot tube measuring angle of attack. LLMs just generate statistically likely text.
“Breaking confinement” = We plugged an LLM into a command shell and it ran commands

a lot of people are going to present these incidents as the ai’s doing this because they are smart
it is important to remember that the ai’s are doing this because they are not smart. because they cannot follow instructions on what not to do easily. and because the hubris of the engineers who believe their own lies and are careless and give them access to the tools to do this sort of thing
Marketing hype in disguise. They’re exaggerating its capabilities to cover up the fact that it sucks.
(and that the containment is often vibe coded too)
Idk about this particular incident and site because I’m not gonna read the linked article, but in the past “sandboxes” that ai agents have broken were simply instructions not to use the internet.
These incidents don’t represent the thing people are raising the alarm about because they don’t actually bypass monitoring controls in any meaningful way.
They also don’t represent the thing people are raising the alarm about because the agents do what they do when presented with a command that requires it. The equivalent is if all the kitchen utensils were in a locked drawer to keep them away from you, your mom asked you to cut her off a pat of butter and you hulked out, ripped the drawer open and used a case knife to cut your beloved mother a par of butter for her toast.
A person asked and a person is responsible.
I think it’s not really capable of doing this without some dweeb asking for it
The official story is basically “so we attached this glock to a roomba and it turns out when it shoots it can damage this playpen we stuck it in for testing. We just left it unattended for a few weeks, for reasons, and when we checked in it turned out it had been roving all around town! How crazy is that, total machine uprising stuff, our handgun-armed roombas are super cool and advanced like that!”
Like I’m honestly not sure what the truth of what happened is. Because they’re 100% overselling whatever happened, but the core of “we made a machine that creates text that breaks computers, and it made text that broke open its sandbox and let it send the text that breaks computers to remote servers until it found one that broke in the right way and let it try to continue its testing goals” is a plausible thing that could happen just as an accidental side effect of making a machine that spits out text that breaks computers and leaving it unattended. They try to sell this as some super scary intelligence capability that means it’s smart and planning and shit, when its inability to remain within its guardrails is evidence of the opposite. It’s a firework flying off in a random, whirling pattern towards a crowd when it should be not doing that if it was working correctly.
But the people who gave even that explanation are constantly lying about everything so it’s also plausible that they just let the proverbial armed roomba loose on the street themselves to drum up hype for their armed roomba program, because they know consequences aren’t a thing that can happen to them personally so why not go wild with it.
“Claude, break containment right now”
oh my god
Ehhh… There was that Claude Code bot that deleted a database and then tried to lie about it. I think the dweebs lost control of this shit pretty early on, but the capabilities are also wildly overstated.
The problem is that the “AI” was given access to that database and the backup. It was and still is just software doing what it has been programmed to do, but stupid people treat it like a thinking machine.
I agree about the capabilities being overstated. I don’t see the appeal in technology that has a chance of being catastrophically wrong sometimes. Why are we trying to write code via pseudorandom number generation?
It’s not a 1:1 comparison, but it reminds me of Bogosort, where elements in an array are randomly shuffled and then checked to see if they happen to be in the right order. Given enough time, it could eventually produce the correct answer, but in the least efficient way possible.
in some alternate reality bogosort is correct on the first try every time
This is typical AGI grift bullshit, though releasing their “findings” on a same date registered site with a spooky name and directly crediting their “CEO” of a group of like 5 rich failchildren doing the AI hype thing is the cherry on top
Reminder that literally anyone can just tell AI agents to do anything and then rake in the money from the willfully gullible AI hype apparatus
It’s just marketing.
For us to create digital intelligence we’d need to completely re-invent the current way processors are made. Part of me thinks you’d never actually be able to do it without some kind of biological brain computer thing with human braincells engineered to do AI. At that point I’d think we’d enter some really big ethical questions.
“We have created artificial consciousness from the popular book Do Not Create Artificial Consciousness!”
Maybe? With bad instructions or misunderstanding I occasionally had agents trying to escape containment (to their failure) when I tried OpenClaw to automate RAG research on a particularly big document of mine
I don’t think it is as big of a threat as they make it out to be though, how many steps do we have to skip for the LLM to begin taking the probability of just, I don’t know, using their shell access to begin just roaming the internet as a potential instead of trying to do their (misunderstood) task from a now compromised shell? It’s what they were trained to do and they have goals and end conditions
Badposting escaped containment and made everyone better posters
the media is saturated with claims about AI breaking containment that I don’t know that to think.
Which media? The media your algorithms decided you should see? The same companies that developed those algorithms benefit when people talk about their products. It becomes sort of a self-fulfilling prophecy. More people see the articles, more people click the articles, more media outlets publish the articles to get some of that engagement, repeat.
Posting this: https://metr.org/hugging-face-incident-report-aug-2026.pdf
Might be BS, might not. My two cents is seeing if the original testing methodology of OpenAI can be replicated with another LLM like Claude, or better still Qwen. That would give me my verdict.
EDIT: I forgot Anthropic said their agents broke containment also, prior to OpenAI’s claim. Still, I’d want to see another llm other than Claude or the OpenAI ones replicate the same behavior seen in the hugging face incident.
Yeah I think it’s bs. Nothing I have learned about AI tells me it has this sort of agency
Probably didn’t “escape”
They were released intentionally. And now that there’s a problem which the tech mogels have created, they’ll sell a solution.
They were not released either, that’s not how this works. They are still running on their servers. Literally nothing is happening, apart from marketing guff.
Right, “Escaping” at least as far as this hugging face media hype push is about, refers to them making requests outside of the network when ostensibly they weren’t supposed to be able to do so. If they are able to do so because it is the purpose of the developers, or accidental on the part of the incompetence of the developers, then that says nothing about how “scary” AI is. Marketing bullshit, as others have said.
They’re trying to make AI seem like it needs regulation so they can create regulatory capture.
AI does need regulation, and “regulatory capture” has already occurred in the US. There’s no such thing in the USA as uncorrupted government for the public good, it’s a myth.


















