The linked article is about some site that OpenAI’s models used to discuss answers and other stuff.
Initially, I believed that Anthropic’s model escaping its sandbox story to be dubious—and OpenAI’s similar story even more so specially since it happened so close to Anthropic’s. I believe that these are just fabrications, more or less, to hype themselves up for cmtheir incoming IPO but the media is saturated with claims about AI breaking containment that I don’t know that to think.
Also, the models I’ve had the chance to use were all free—which were good for non-trivial but repetitive tasks but not much else—so I don’t know the capabilities of the flagships.



I think it’s not really capable of doing this without some dweeb asking for it
The official story is basically “so we attached this glock to a roomba and it turns out when it shoots it can damage this playpen we stuck it in for testing. We just left it unattended for a few weeks, for reasons, and when we checked in it turned out it had been roving all around town! How crazy is that, total machine uprising stuff, our handgun-armed roombas are super cool and advanced like that!”
Like I’m honestly not sure what the truth of what happened is. Because they’re 100% overselling whatever happened, but the core of “we made a machine that creates text that breaks computers, and it made text that broke open its sandbox and let it send the text that breaks computers to remote servers until it found one that broke in the right way and let it try to continue its testing goals” is a plausible thing that could happen just as an accidental side effect of making a machine that spits out text that breaks computers and leaving it unattended. They try to sell this as some super scary intelligence capability that means it’s smart and planning and shit, when its inability to remain within its guardrails is evidence of the opposite. It’s a firework flying off in a random, whirling pattern towards a crowd when it should be not doing that if it was working correctly.
But the people who gave even that explanation are constantly lying about everything so it’s also plausible that they just let the proverbial armed roomba loose on the street themselves to drum up hype for their armed roomba program, because they know consequences aren’t a thing that can happen to them personally so why not go wild with it.
“Claude, break containment right now”
oh my god
Ehhh… There was that Claude Code bot that deleted a database and then tried to lie about it. I think the dweebs lost control of this shit pretty early on, but the capabilities are also wildly overstated.
The problem is that the “AI” was given access to that database and the backup. It was and still is just software doing what it has been programmed to do, but stupid people treat it like a thinking machine.
I agree about the capabilities being overstated. I don’t see the appeal in technology that has a chance of being catastrophically wrong sometimes. Why are we trying to write code via pseudorandom number generation?
It’s not a 1:1 comparison, but it reminds me of Bogosort, where elements in an array are randomly shuffled and then checked to see if they happen to be in the right order. Given enough time, it could eventually produce the correct answer, but in the least efficient way possible.
in some alternate reality bogosort is correct on the first try every time