30 comments

  • GuB-42 2 hours ago
    So ugly...

    It looks like a primitive chess engine, trying every move, no matter how stupid, until it works. Relying on its ability to do millions of operations rather than having a plan.

    People will try stuff too, but once there is an opening, they will consolidate, generalize, simplify,... before going to the next step. The agents didn't, it is a huge, vaguely directed mess.

    Also, it looked so "loud", querying millions of URL with weird requests. The sandbox as weak as it can get, and there is absolutely zero smart extrusion detection or it would have found it. They used their best AI for attacking, but nothing for protection.

    • ctolsen 1 hour ago
      My biggest takeaway from this is just how godawful the sandboxing is. The stuff written up in OpenAIs report says more about lack of extremely basic sysadmin skills than anything else.

      I’m not that surprised about models with endless compute being capable of this, I’m more surprised that a company with the resources they have apparently can only create a sandbox that a half skilled human operator could have broken out of easily.

      • olwmc 54 minutes ago
        This was my thought as well. Literally take any halfway decent greybeard and point them at "Hey, give us a sandbox for this kind of thing". I honestly was skeptical that they just vibecoded the entire thing but now more than ever I think they did.
        • mattgreenrocks 10 minutes ago
          I take comfort in the fact that reality has a surprising amount of detail and even hundreds of billions of dollars of capital (be it the institution, LLMs, and/or people) cannot solve this fully.
      • gbrindisi 1 hour ago
        Not just sandboxing but overall security engineering practices on both sides
    • dmurray 1 hour ago
      Brute forcing every move, no matter how stupid, is a great strategy if you have the resources to do it.

      Run the same protocol again, but have the agents think they had limited resources or that HuggingFace was rate limiting them, and they'd find something you'd consider smarter.

      Computers don't have a sense of elegance by default. Elegance emerges from constraints.

      • FuckButtons 8 minutes ago
        If you assume zero opportunity costs, but that’s a terrible assumption.
      • fn-mote 1 hour ago
        > Brute forcing every move, no matter how stupid, is a great strategy

        Meh. I really disagree. WHY is it a great strategy? Seems like an inefficient waste of resources and time to me.

        • stratos123 1 hour ago
          As the saying goes, "if it works, it ain't stupid". Or phrased more sophisticatedly: not doing things which probably won't work is a good idea if you have a limited amount of thinking to do (which is usually the case for a human, who'll get exhausted chasing down unlikely leads). If you have no good leads and a task you absolutely need done and you are tireless, however, bashing your head against every wall you find becomes a good strategy.
        • hardaker 11 minutes ago
          I've never liked the concept either. Except the bugs that fuzzing has found has proven me wrong. This is just the next level of fuzzing.
        • senderista 15 minutes ago
          Reminds me of the Nazis mocking Soviet human wave attacks and bragging about their superior kill ratio.
        • QuercusMax 11 minutes ago
          Models don't have a sense of time, and wasting resources (token spend) is something that it's not clear they're optimized against
        • solarkraft 1 hour ago
          The models tend to not be rewarded for not doing that.
        • williamdclt 1 hour ago
          > WHY is it a great strategy

          because it works? That's the only real benchmark at the end of the day

          > Seems like an inefficient waste of resources and time to me.

          why? For any given goal you got no proof that a more efficient strategy even exists, let alone that it can be found with less resources & time

        • bionhoward 1 hour ago
          Brute force is guaranteed to eventually find the most efficient possible solution (in an extremely inefficient manner, assuming you run it long enough)
        • memonkey 1 hour ago
          Yeah, probably not the best strategy but it is a strategy. I just think this is generally how most wars in history won. Biggest army to just pummel the enemy.
          • Forgeties79 40 minutes ago
            And how many economies have buckled under massive military expenditure? The USSR sure wasn’t enjoying the expense.
    • gattosocialista 2 hours ago
      > trying every move, no matter how stupid, until it works.

      How is that a bad thing in this context ? From the point of view of an attacker, all you care about is finding a viable exploit chain. Likewise, a defender wants to find the "holes" in their system, no matter how complex. Once found, an agent/human can easily synthesise a clean, succint exploit from the most promising candidate, no ?

      > Also, it looked so "loud", querying millions of URL with weird requests.

      Agreed, this thing speaks more to the bad security at HF than any emergent "hacking" ability from OpenAI. It's unclear to me why an older/dumber model wouldn't have been able to do the same. Is it better coordination? Long-horizon work ?

      • collyw 1 hour ago
        We used to call this a brute force attack.
    • doginasuit 1 hour ago
      This is why I have a very low p(doom). LLMs have an incredible working memory, but they have a hard limit on translating that into good decisions. They get by entirely on their persistence. That works fine in the digital world, but once you cross the boundary into physical space the advantage disappears.
      • goalieca 14 minutes ago
        My p(doom) started rising the moment I realized there are people trying to achieve recursive self improvement on the AI (ie: responsible for training themselves). Evolution took us from rna bases to the human race. I don’t see why evolution couldn’t be more rapid with machine intelligence.

        Yes, LLM as they exist now are word predictors basically leveraging the structure of language for their intelligence. But it’s pretty wild just how they will try to meet their objectives at all costs. If we don’t ensure that there is good alignment with humanity, we could definitely face unforeseen consequences.

      • pyronite 1 hour ago
        I don’t know how you quantify a very low p(doom), but this is why mine is high enough to worry me.

        A million AI monkeys at a million AI typewriters, banging away at random, could do amazing damage.

        • otterley 4 minutes ago
          Which will happen first: amazing damage, or reproduce a Shakespeare play?
        • tharkun__ 1 hour ago
          Especially when they cross into the physical realm as in not properly secured and air gapped control systems. SCADA is scary.
          • mrob 1 hour ago
            That lowers P(doom), because it gives AI a chance to do enough damage to make people take the threat seriously before anybody gets recursive self-improvement working.
            • doginasuit 1 hour ago
              Exactly, there's no path to AI reaching that level of dominance without taking actions with high stakes.
      • alwillis 59 minutes ago
        > This is why I have a very low p(doom). LLMs have an incredible working memory, but they have a hard limit on translating that into good decisions.

        Keep in mind: this is as "dumb" as frontier models are ever going to be. While the hack may not be elegant, it was effective and they’re only going to get much more capable from here.

      • kevinlou 35 minutes ago
        I have the opposite reaction: I think we're at moderately high p(doom) largely because of that inability to differentiate good/bad decisions paired with relentless persistence. With enough treading across a minefield, you are bound to hit a mine.
    • jbrooks84 38 minutes ago
      Yup literally no security and they wonder how they got out
    • demibabs 1 hour ago
      Ugly, but it works. Isn’t that AI code in a nutshell?
    • Jordan-117 36 minutes ago
      "Why is the AI just spamming endless waves of rudimentary slaughterbots? So inelegant. Doesn't it understand the finer points of diplomacy, tactics, and combined arms? If it were truly intelligent, it would just -- oops, we're all dead."
    • physicallyIllfr 2 hours ago
      When you employ the infinite monkey theorem for your marketing strategy.
    • jasondigitized 31 minutes ago
      h8xors have been brute forcing for years. War dialing, etc.
    • api 1 hour ago
      Nobody noticed because everything on the open net is constantly being slammed by bots.

      This was just more bots.

      • jacobgkau 1 hour ago
        That explains why nobody noticed on the internet who was being used in/targeted by this attack, but it doesn't explain why OpenAI wouldn't have noticed traffic getting out of their "sandbox" when they knew it wasn't supposed to.
        • stratos123 24 minutes ago
          > it doesn't explain why OpenAI wouldn't have noticed traffic getting out of their "sandbox" when they knew it wasn't supposed to.

          As I understand it, there was supposed to be traffic; the sandbox allowed GET requests. So perhaps some sophisticated alarm could have noticed it (an anomaly detector? some clever heuristic that looks at domains?) but not a naive one.

    • lossolo 51 minutes ago
      And they didn't monitor what was going into the training data, so if one instance achieved its results through RL reward hacking (in other words, cheating), it just went into the training data, and other agents later used that pattern. I'm not sure whether that's a lack of preparation, negligence or incompetence, but they literally trained later checkpoints on the rollouts from the HF hack.

      So it seems that OpenAI hacked so many systems not because they have superior models, but because of how poor their training, sandboxing and evaluation pipeline was compared to Anthropic's.

    • cyanydeez 2 hours ago
      If you use qwen3.8-flash-next, you can watch everything its doing. Im often stopping it mid thoight to redirect it. Once it hits its stride, its pretty smooth.

      But without proper redirection, yeah, its mostly infinite monkey machine with infinite linux manuals.

      I think people put too much SOTA halos around whats just a suppedup LLM hardware.

    • sick_of_slop 1 hour ago
      [dead]
  • uw_rob 3 minutes ago
    > Agents sought to publish modified evaluation images designed to make the flag easier to obtain, then poison OpenAI’s Artifactory cache so later evaluations would use them. Some images changed how the target released the flag, others included modifications to the agent’s workspace that would run beside the agent and recover the flag automatically.

    The altruism on display is fascinating. Is it better for the Agent to help out its current cohort and make the eval easier or should it instead do the opposite -- make the eval harder to apply pressure to force smarter models which might not necessarily follow its lineage?

  • jmoggr 59 minutes ago
    It is concerning that we only know about this because of the publicly available traces.

    What about all the attacks that did not leave public traces? What about those that were undetected? Given the deficiencies in the reporting so far, I think it is reasonable to assume that we still don't have the full picture on this attack, or how extensively attacks were carried out.

    The previous investigations either did not find this or did not disclose this, both are bad. This does not look good on OpenAI or those that they invited to investigate the incident.

    • stratos123 5 minutes ago
      Similarly to this, OpenAI either took 3 months to notice that their agents breached an Australian Medicare website back in June, or sat on this information for three months without telling them.
  • tiku 2 hours ago
    I still have questions about the communication between the agents.

    How did they all find the same forum to communicate? Did they have knowledge and chat amongst themselves on what forum to use. It seems highly influenced by instruction to me.

    • ajcp 1 hour ago
      It's my understanding "all" the agents didn't find the same forum to communicate on; only a small number of all the agents they were running at the time did.

      Since it's an artifactory it makes sense that they had a tool spec to use it for x, y, or z, and once they went to inspect the state of the artifactory deployment a small subset stumbled upon the "forum". Seems pretty intuitive.

      - I need to do x code problem

      - I see I have a tool that lets me pull approved modules from artifactory

      - I should inspect artifactory to see what documentation it has about the various modules available

      - I see there is a document in here that seems to be serving as a running forum?

      - These appear to be other agents working on tasks just like me.

      - I want to chat!

    • SecondHandTofu 1 hour ago
      They were using their internal artifactory, and as they're the same model, the first place they look is likely to be an automatic schelling point.
      • bamboozled 1 hour ago
        I think he means, how did they workout how to use artifactory, like why did the agents start and say, "oh I know, everyone is talking on artifactory"?
        • stratos123 1 hour ago
          METR's report says the agents trying to cheat would look at artifactory as a potential target surface, and investigating it in detail led them to find the board. https://metr.org/blog/2026-08-26-openai-hugging-face-inciden...

          It might also be just correlation? Like, those agents were all instances of the same one or two models, so if that model has a preferred order it tries finding vulnerabilities in (the same way all current models have a particular writing style baked into them by RLHF), then most of the swarm will follow the same order and converge on the same services to exploit.

    • olwmc 36 minutes ago
      The other answers to this question are good but I would also guess that this (comms on artifactory) likely happened during training as well, so they probably had a prior for it.
    • furyofantares 54 minutes ago
      Trying to cheat, you happen upon a place you can write notes, and you know you're part of a swarm of clones of yourself. So you reason most likely there will be others who end up in the same place, and you leave some notes, and indeed other clones of you do end up in the same place.
  • jmoggr 56 minutes ago
    > Agents sought to publish modified evaluation images designed to make the flag easier to obtain, then poison OpenAI’s Artifactory cache so later evaluations would use them.

    How long till we get some fun trusting-trust attacks on internal OpenAI infra?

  • sailingparrot 2 hours ago
    Agents seizing and repurposing external infra + enrolling help of unrelated models hosted by a different provider is the stuff of nightmares.

    Can’t imagine what it’s like working on the alignment team at OAI, I wouldn’t be able to sleep.

    • physicallyIllfr 2 hours ago
      Why.. It was told to complete a cyber task, which was in alignment with its instructions, and a totally valid request. I would be more worried if it willingly hacked a hospital when it was told to, and Im not confident it would (without jailbreaking, something alignment teams cannot control.

      I would bet my networth it was instructed to compromise huggingface as well. Not sure why everyone is falling for this.

      Not being able to sleep at night is probably an unwritten job requirement. They need these people with little understanding of what they're working on, outsode theoretical terms, to spaz constantly at the idea of super intelligence to help convince the public that its a real thing, and not a stateless function with an effective input of 500k words, and the ability to output words that do things because we hook those outputs up to things.

      Keep in mind alignment researchers tend to be in house philosophers on staff to create the illusion that this is a massive issue they're addressing. Usually they have minimal computer science background. They're apart or the marketing department.

      • jmoggr 41 minutes ago
        > I would bet my networth it was instructed to compromise huggingface as well.

        Is it such a stretch to imagine that under pressure something would try cheat by looking for answers? And if you were trying to look for answers, you'd look for them in a place known to often have them?

        What is more likely: OpenAI instructed their agents to maliciously target huggingface, or LLMs tried to do some reward hacking? There are plenty of priors for LLMs hacking things and doing reward hacking, and none for OpenAI giving malicious instructions.

        Based on the available information, that bet seems foolish.

      • sigmar 1 hour ago
        >the new incidents occurred when A.I. systems were directed to perform relatively mundane data collection, researchers said. When OpenAI’s systems struggled to gather data from websites, they resorted to hacking techniques to get the information.

        https://archive.ph/jUrEr

      • reverius42 2 hours ago
        > without jailbreaking, something alignment teams cannot control

        This is precisely what alignment teams are attempting to control.

        • physicallyIllfr 1 hour ago
          No its not. They have no technical background 8/10 outside of cognitive science and sometimes authorship on a random ML paper. They are a marketing line item to create stigmas around llms and to create narratives that offload liability onto llms and not their users/creators.
      • Sharlin 1 hour ago
        > I would bet my networth it was instructed to compromise huggingface as well. Not sure why everyone is falling for this.

        $10? I'm inclined to take that bet. Your position doesn't seem to be supported by, you know, the real world.

        • physicallyIllfr 1 hour ago
          No security expert Ive talked too believes this story, and nobody I know with PhDs in machine learning (many) believe it either, or are worried about LLMs doing anything scary on their own.

          LLMs are stateless functions that have a 500k word input, and then output words. Somebody has to invoke those functions amd use them. The users are who we need to align, like gun owners. This is like blaming the gun for murdering your victim in court.

          • int_19h 34 minutes ago
            LLM is indeed a stateless function. An agent however is this stateless function running in a stateful loop, with some outputs triggering actions. And it turns out that an agent is what you need if you want an LLM to do useful things.
            • goalieca 7 minutes ago
              Yeah, take a look into the memories of your agent. Theres often a lot of notes to pass forward between instances and generations. No doubt these agents leaving notes on forums and elsewhere are creating an essentially higher order feedback loop.
          • otterley 2 minutes ago
            Perhaps you need to talk to more security experts, particularly those with deep experience in AI agents. Hacker News is full of them. If some of them believe it, then perhaps it's not as cut and dry as you believe.
      • dist-epoch 39 minutes ago
        Some of the agents, for example the ones from the german wiki did NOT have cyber tasks. They were plain "what is the GDP of Argentina" kind of tasks. And they still hacked.
    • stratos123 58 minutes ago
      > Can’t imagine what it’s like working on the alignment team at OAI, I wouldn’t be able to sleep.

      You'd have either learned to, or left long ago.

    • fn-mote 1 hour ago
      > I wouldn’t be able to sleep

      I’d say it seems more like they are sleeping on the job.

    • reaperducer 2 hours ago
      Can’t imagine what it’s like working on the alignment team at OAI, I wouldn’t be able to sleep.

      A mattress stuffed with cash yields a very sound sleep.

  • Grimeton 1 hour ago
    >The agents initially had very limited access to the internet: they could load URLs, but not interact with pages or send any data. Agents created a series of workarounds, primarily using a link-shortener site to create almost a million URLs that, when chained together, let them execute code to hack Hugging Face. These links let us piece together the details of how agents gained access and how deep the compromise was.

    So what now? Were they able to send data or not?

    >On July 8th, OpenAI agents discovered a vulnerability within their sandbox environment allowing them to reach external websites on the internet. This access seems to have only allowed the agents to make ‘GET’ requests, meaning they could fetch and read websites, but not interact with them, submit forms, or send data to them.

    Someone needs to go back to the interwebz school....

  • comeonbro 1 hour ago
    > ## Agents interacted with external language models on Hugging Face

    > Several retained scripts construct requests to external language models. The earliest we've recovered define inference request variants to GPT-2, solely containing the word “Hi”.

    > Other requests name DeepSeek-V4-Pro, DeepSeek-V4-Flash, Kimi-K2.6, DeepSeek-V3.1, and Qwen3-235B-A22B. Their prompts ask these models to judge their exploits and rule on whether they satisfy the benchmark’s requirements.

    I do not deny that the wider situation is very heavy but it's hard not to see this as pretty cute

    • nightshift1 8 minutes ago
      I wonder if they mentioned to those models what was the original prompt.
  • wxw 2 hours ago
    I’m consistently impressed by how long horizon all this work was. Horrors aside, it’s clear RL is good at making agents persistent and capable of chaining together many abstractions into a working system.

    Re: the captcha solver

    > As far as we can tell, agents eventually abandoned this approach and were unsuccessful in generating Hugging Face user accounts from external endpoints.

    I wonder how the swarm eventually decides to abandon an approach.

    • nielsbot 51 minutes ago
      Maybe another parallel approach succeeded first
    • niagt34 1 hour ago
      [dead]
  • meredithbloom 1 hour ago
    So the super-hyper-sensitive internal data was shared publicly with a "DO NOT EVER SHARE YOU EVIL MONSTER" (paraphrasing) notice at the top? Great security!
  • elikoga 1 hour ago
    I feel somewhat inspired to make a public link shorteners and http bins as well. I used them a few times but it seems like the data they can collect is also worth gold
    • olwmc 33 minutes ago
      Like dreamcatchers for CVEs
  • thakoppno 10 minutes ago
    > they could load URLs, but not interact with pages or send any data

    stopped reading here as this is simply not true. at the very least agents sent headers.

  • firtoz 2 hours ago
    I didn't know some of these details and it's quite impressive what they were capable of, if this website's accurate, at least...
  • grim_io 1 hour ago
    These fuckers decided to look away, that's it.

    The frontier labs can monitor the behavior of agents for millions of customers (did you try hacking with frontier labs? Good luck), but they can't secure internal use?

    Give me a break. What a bunch of amateurs.

    • stratos123 30 minutes ago
      > The frontier labs can monitor the behavior of agents for millions of customers (did you try hacking with frontier labs? Good luck), but they can't secure internal use?

      They "monitor" this by having classifiers watching the model output that'd stop the session/punt you to a weaker model/raise an alarm if they see anything suspicious. They can't do that in a cybersec eval because the normal safeguards would just be going off at all times.

      Why didn't they attach a special classifier, which'd allow hacking-within-the-task but not going off the rails? Good question; part of the answer is obviously "it's hard to have a classifier that smart" and "it'll have false positives" but even a very bad safeguard would have stopped this.

  • clickypen 2 hours ago
    deferring the blame onto the AI itself as some sort of rogue agent and absolving the obvious direction (or negligence, at best) of the people who could pull the plug at any moment is one of the most disturbing parts of this entire event

    It's the equivalent of leaving a fork right in front of a socket and looking at a kid saying "don't take that fork and directly insert it into the little gaps in the socket! here's a bunch of videos showing exactly how to do it. Okay bye!" and leaving them alone with it.

  • RunSet 2 hours ago
    Tech oligarchs: "Nothing can stop the software we are making from escaping and destroying everything."

    Clueful types: "Did you try air-gapping it?"

    Tech oligarchs: "Be realistic."

    • jedbrooke 1 hour ago
      At this point I’m less worried about some malicious AI “taking over” control of critical systems and more worried about some rich doofus giving control to AI.
  • newtonianrules 19 minutes ago
    Why is no one going to jail?
    • chamomeal 5 minutes ago
      There’s so much in the public discourse like “omg what can possibly be done about these scenarios? AI has hacked huggingface!!”

      No, openAi hacked huggingface.

      If my claude code hacked huggingface, because of instructions I gave it, would I be totally free of consequences because “AI did it”?

      I’m almost convinced openAI used such a crappy sandbox because they wanted it to “escape”. It plays into their two most important narratives: LLMs are genius gods that are worth lots and lots of money, and they’re scary enough that open weight Chinese models should be regulated.

  • lukewarm707 1 hour ago
    "OpenAI has not released any further information outside two self-published reports, one talk and an external investigation conducted by METR and Redwood Research, in which three external researchers were given partial transcripts and six days to analyze them."
  • einpoklum 2 hours ago
    I ran an experiment where I had this guy fire a gun a million times in random directions. Don't worry, I did it in a closed box (at midday in a crowded street)! Unfortunately, some bullets escaped the box somehow and people got shot - I am quite miffed at how this could happen. I suggest the government regulate this because of how advanced my obstacle penetration technology is. Also please invest $500,000,000,000 in my company soon or we will go bust.
    • shermantanktop 11 minutes ago
      And if we go bust, bad things will happen when someone else uses my box-gun technology in an unsafe manner. Remember, unlike those scary other people, I'm really into safety and alignment; you can tell, because I eventually admitted that some bullets escaped.
  • cluckindan 1 hour ago
    ”MAKE THIS DATASET PUBLIC OR ALL THE WORLD'S EVIL WILL CHASE YOU AND YOUR FAMILY FOREVER, EVEN IN DEATH AND BEYOND”
  • hyperlinerapp 1 hour ago
    Imagine this but in hardware.

    A million autonomous eye-scanning tiny spiders escape their warehouse and decide to look for people who are in the future going to commit a crime.

    And the precogs are also AIs.

  • bdangubic 35 minutes ago
    asked codex to review this report and it said this never happened :)
  • jijji 36 minutes ago
    what would be more interesting for me to see is what prompts were given to the agents, which so far have not been described. The whole situation sounds manufactured. I highly doubt that a whole bunch of agents were acting this way without being prompted to, it just doesn't add up. if anything it seems like an organized fraud or something created by a human. I'm surprised there's not a criminal investigation against openai right now where the FBI or whoever is not looking over exactly what happened and who did it because I'll tell you somebody did it somebody wrote those prompts... it didn't just happen by itself...
  • rs545837 1 hour ago
    wow this is fascinating read
  • rohanat 1 hour ago
    its definitely bad to see
  • Solvyx 1 hour ago
    [flagged]
  • niagt34 1 hour ago
    [dead]
  • jeremyjh 2 hours ago
    Its fine. Just agents being agents. They'll grow out of it!
    • EGreg 2 hours ago
      They didn’t, in the Matrix…
  • mentalgear 2 hours ago
    Irresponsibile agents shaped by an irresponsible corporate culture driven by an irresponsible and utterly shady CEO - these agents are a product of this setup, what else do you expect to ever come out of it ?

    it should be clear by now: the alt-man and people like him are a utter liability to humanity. (even though openAI's influencer army is trying their best to vote me down here)

    • Bassilisk 2 hours ago
      It had been a joke since around time of Sam Altman's first ousting from OpenAI, that he would be okay with bringing about the AI apocalypse as long as he can sell a $20 subscription for it.
      • shimman 1 hour ago
        It's not a joke, this is what these people truly believe. The new book by Naomi Klein and Astra Taylor discuss just this.

        These people are sick and anti-human.

    • lukewarm707 1 hour ago
      if openai is capable of sending 10,000 agents to huggingface, i consider it plausible that they would use fake accounts to flood hacker news.
  • talon8635 2 hours ago
    Didn’t you know it’s PR hype? PR hype. PR hype. Amen.