The linked article is about some site that OpenAI’s models used to discuss answers and other stuff.

Initially, I believed that Anthropic’s model escaping its sandbox story to be dubious—and OpenAI’s similar story even more so specially since it happened so close to Anthropic’s. I believe that these are just fabrications, more or less, to hype themselves up for cmtheir incoming IPO but the media is saturated with claims about AI breaking containment that I don’t know that to think.

Also, the models I’ve had the chance to use were all free—which were good for non-trivial but repetitive tasks but not much else—so I don’t know the capabilities of the flagships.

  • OptimusSubprime [he/him, they/them]@hexbear.net
    link
    fedilink
    English
    arrow-up
    6
    ·
    edit-2
    2 days ago

    Posting this: https://metr.org/hugging-face-incident-report-aug-2026.pdf

    Might be BS, might not. My two cents is seeing if the original testing methodology of OpenAI can be replicated with another LLM like Claude, or better still Qwen. That would give me my verdict.

    EDIT: I forgot Anthropic said their agents broke containment also, prior to OpenAI’s claim. Still, I’d want to see another llm other than Claude or the OpenAI ones replicate the same behavior seen in the hugging face incident.