The linked article is about some site that OpenAI’s models used to discuss answers and other stuff.
Initially, I believed that Anthropic’s model escaping its sandbox story to be dubious—and OpenAI’s similar story even more so specially since it happened so close to Anthropic’s. I believe that these are just fabrications, more or less, to hype themselves up for cmtheir incoming IPO but the media is saturated with claims about AI breaking containment that I don’t know that to think.
Also, the models I’ve had the chance to use were all free—which were good for non-trivial but repetitive tasks but not much else—so I don’t know the capabilities of the flagships.



Idk about this particular incident and site because I’m not gonna read the linked article, but in the past “sandboxes” that ai agents have broken were simply instructions not to use the internet.
These incidents don’t represent the thing people are raising the alarm about because they don’t actually bypass monitoring controls in any meaningful way.
They also don’t represent the thing people are raising the alarm about because the agents do what they do when presented with a command that requires it. The equivalent is if all the kitchen utensils were in a locked drawer to keep them away from you, your mom asked you to cut her off a pat of butter and you hulked out, ripped the drawer open and used a case knife to cut your beloved mother a par of butter for her toast.
A person asked and a person is responsible.