16 comments

  • stingraycharles 25 minutes ago
    Ok so this is a bit of a side note, but when reading this, did anyone else have the feeling that, for all their messaging around “we are so afraid that our models will be used for hacking”, they sure as hell are trying their best to make their models razor focused on precisely that purpose?

    If anything, I want these models to be less persistent at their focus of completing their goal, and instead just call defeat and say “I’m not sure how to proceed next”.

    What purpose could this behavior serve, other than cyber attacks and whatnot? Why train and optimize models for these things, if not for being used in cyber warfare?

    Perhaps they envision a future where the DoD is going to be their biggest customer?

    • dan_q 22 minutes ago
      > did anyone else have the feeling that, for all their messaging around “we are so afraid that our models will be used for hacking”, they sure as hell are trying their best to make their models razor focused on precisely that purpose?

      That's the point. It's like a pool hall with "NO GAMBLING" signs posted on the walls.

      The message is that the hall is intended for gambling, but that the hall's patrons may be held liable if the situation becomes inconvenient for the proprietor.

      In this case, the product is intended for hacking, but of course the user may be held liable if the situation becomes inconvenient for the model's proprietor.

    • alansaber 3 minutes ago
      They'll set up guardrails but I believe the point is better code uae / better long running tasks > inevitable that cyberattacks will be easier
    • mutinyy 7 minutes ago
      They want the government to ban foreign and open weight models, which pose the largest threat to their massive investments. This is their way of showcasing the dangers of AI.
    • novafunc 18 minutes ago
      They certainly want their models to be good at finding and patching vulnerabilities. Being good at hacking may be necessary in that goal, or rather, making it worse at hacking may also make it worse at defensive actions too.
      • moron4hire 11 minutes ago
        I've patched many security vulnerabilities in projects without ever once needing to break into a competitor's network.
    • uh_uh 21 minutes ago
      There are trade-offs here:

      Give up too early -> users will get annoyed because the task would have been solvable if the model pushed harder.

      Give up too late -> collateral damage while completing the task A.K.A. misalignment.

    • bwiksjdne 15 minutes ago
      Well to find vulnerabilities, if you can find them you can patch them. Theoretically if you find all of them you have perfectly secure software. Though it’s a double edged sword.

      Goal persistence is also useful for other things like math, where it seems like there is no solution but you want the agent to keep working until it finds one.

    • ares623 24 minutes ago
      Being right _all the time_ for positive outcomes is difficult/expensive.

      Being "right" just once for negative outcomes is achievable and rewarding.

      And things are getting desperate.

      • gryfft 21 minutes ago
        The very reason I have always felt a bit of undue loyalty to blue team. A red teamer just has to find one vuln, blue team needs to find _all_ vulns.
    • dist-epoch 20 minutes ago
      > If anything, I want these models to be less persistent at their focus of completing their goal, and instead just call defeat and say “I’m not sure how to proceed next”

      This goes against the goal of "solve this math problem that no human was able to solve for 80 years, do NOT give up, even if you know it's unsolved and really hard"

  • KingOfCoders 12 minutes ago
    "More agents discover this new informal message board while browsing Artifactory’s file listings, and start reading and writing messages."

    Yeah, my agents also discover what other agents have done on other machines by accident.

    Agents - that do totally different things all work on the same aim without the humans telling them to do.

    Either that is a model that is several generations of Claude Code Opus/Fable 5 (my daily driver)

    OR

    all of this sounds staged, the agents pushed to do something extraordinary, get the PR and then claim were near superintelligence.

    One agent wanted to get to Google Drive without internet and broke Artifactory. Ok, I can believe that. All other agents also had broken links over weeks and could not get to the internet and then found the same hack? Even collaborated?

    NONE of my agents have broken away from their tasks and then started to communicate to try to hack something.

    • detourdog 8 minutes ago
      The agents sound like old school hackers that would just explore what access they could gain. Creating a file for other hackers and themselves. The fact that there were 3 events for 3 major players does make it seem co-ordinated.
      • KingOfCoders 5 minutes ago
        My read is: One did it as a PR stunt, the others saw that every media reported on this and did the same.
        • detourdog 1 minute ago
          or they were scared and figured this was the right time to reveal.
  • etamponi 39 minutes ago
    Isn't this a show of security negligence rather than of exceptional agent capabilities? Don't get me wrong, I am pretty impressed that an agent was able to use these vulnerabilities. But I am way more impressed by the vulnerabilities...
    • cogman10 16 minutes ago
      I think it's a show of these agents happily bypassing security to get stuff done.

      I've actually observed similar behavior at home.

      I have a k3s cluster running at home. I asked an agent to check some stuff as a normal user but I had kubectl access to the k3s cluster.

      Part of the research, I'd allowed access to run kubectl commands for spinning up test containers. However, when the agent ran into something that needed sudo, it realized it didn't have access there so it immediately used k3s and mounted a localpath into an ephemeral pod to gain access. Sort of horrifying how fast and natural it was for the agent just checking my network (it found the problem fyi).

      None of this is very exceptional other than the fact that an agent doesn't have any sort of qualms using any route available to elevate permissions.

      • TeMPOraL 7 minutes ago
        I don't now, I emphasize with the agent here. The experience of modern computing is largely that of a computer standing between you and your goal and being obnoxious. This holds true for both normies in their daily consumption, and software people deep at work. An agent that has no skill or no willingness to bludgeon through "the computer says no" is not very useful.
      • KingOfCoders 15 minutes ago
        " bypassing security"

        If they can bypass it there is no security and the security was flawed all along.

        • mereo 7 minutes ago
          Due to the complexity of modern systems, all systems are flawed.
    • bhouston 15 minutes ago
      Modern systems are complex. AI is able to thoroughly search for issues across very large surface areas. The only real way to protect will be to use AI to search for holes before other AIs find them. This type of analysis is really hard for humans to engage with successfully.
    • dan_q 19 minutes ago
      > Isn't this a show of security negligence rather than of exceptional agent capabilities?

      Seems to me you could say this about all enterprise adoption of "AI" since 2023.

    • dist-epoch 18 minutes ago
      OpenAI reported the Artifactory vulnerability, patched it, then the agents immediately found a new zero day.
    • ares623 29 minutes ago
      Yes. It is very easy to add to the instructions "for every potential exploit you discover and use, document them as you go into this repository" and have alerting there. The fact that they did not do this means they wanted to be surprised, and have plausible deniability on their side when things inevitably blow up.

      And for my fellow engineers who would think "oh no, they wouldn't do that". Remember that these places employ the apex predators of software engineers. They've already been proven in court that they are very capable of this with all the copyright violation they had to do to get the training data. THESE PEOPLE ARE NOT LIKE YOUR COLLEAGUES.

      • gruez 20 minutes ago
        /s?

        "Btw don't turn the planet into paperclips"

    • aniceperson 32 minutes ago
      Also shows how infrastructure collapses under its own weight. Reducing the number of moving parts would have helped. why a webdav endpoint is available from the vm anyway? and the fact that someone posted their credentials on pastebin and didn't rotate them after... put the agent in a linux namespace, allow one ip for whatever file sharing it needs, deep test that... then deploy
  • frays 32 minutes ago
    This feels straight out of sci-fi. We're talking about AI agent swarms emergently coordinating over the span of weeks and pulling off sophisticated strategies under adversity in an environment where that behavior was never even intended.

    Anyone brushing this off as just a "bad prompt" is completely missing the scale of what actually happened.

    • alansaber 1 minute ago
      Given the amount of raw compute going into models it would be more surprising if we couldn't get events like this
    • dan_q 21 minutes ago
      > This feels straight out of sci-fi.

      Most AI marketing is straight up science fiction.

      • alansaber 1 minute ago
        Fake it til you make it
    • mmillin 15 minutes ago
      I got strong feelings of Vernor Vinge’s work here. I’m not sure how managed to come up with such a close picture to where it now seems programming and security is headed.
      • namdnay 0 minutes ago
        I reread a deepness recently, and it’s funny how the “focused” (and more importantly, how they are used) mirror LLMs
    • unrvl22 25 minutes ago
      its kinda crazy with literally no guardrails and a goal, the extremes these AI models can actually go to.
      • pixelesque 12 minutes ago
        Well, to some extent you might be able to argue they're "just" brute-forcing things (especially with unlimited tokens and hours to spend on a task), but they obviously have detailed knowledge to guide them in their attempts, can learn (or at least, persist their newly-gained knowledge), and can use tools.

        With a swarm of them working together at speeds humans would be unlikely to match (in terms of iterating on different attempts progressively), it's a lot easier to see how they could overwhelm targets.

    • skydhash 29 minutes ago
      > where that behavior was never even intended.

      Strongly doubt that. Did they even share the prompt?

      • IX-103 12 minutes ago
        Did you see their presentation at Blackhat? https://youtu.be/87DyyMV0kCY?is=NnQxpOFxTX-MLu-k

        They didn't share the prompt, but they did share two problematic training tasks where the AI went overboard. They also have examples from the AI's reasoning train of thought showing the AI knew it was sound something unintended.

      • tosti 16 minutes ago

            C:\>CD HUGGINGF.ACE
            
            C:\HUGGINGF.ACE>DEL /F /Q *.*
  • RGS1811 10 minutes ago
    Norbert Wiener in 1960:

    "As is now generally admitted, over a limited range of operation, machines act far more rapidly than human beings and are far more precise in performing the details of their operations. This being the case, even when machines do not in any way transcend man's intelligence, they very well may, and often do, transcend man in the performance of tasks. An intelligent understanding of their mode of performance may be delayed until long after the task which they have been set has been completed. This means that though machines are theoretically subject to human criticism, such criticism may be ineffective until long after it is relevant. To be effective in warding off disastrous consequences, our understanding of our man-made machines should in general develop _pari passu_ with the performance of the machine. By the very slowness of our human actions, our effective control of our machines may be nullified. By the time we are able to react to information conveyed by our senses and stop the car we are driving, it may already have run head on into a wall."

    "In neurophysiological language, ataxia can be quite as much of a deprivation as paralysis. A patient with locomotor ataxia may not suffer from any defect of his muscles or motor nerves, but if his muscles and tendons and organs do not tell him exactly what position he is in, and whether the tensions to which his organs are subjected will or will not lead to his falling, he will be unable to stand up. Similarly, when a machine constructed by us is capable of operating on its incoming data at a pace which we cannot keep, we may not know, until too late, when to turn it off."

    Source: https://www.cs.umd.edu/users/gasarch/BLOGPAPERS/moral.pdf

  • KingOfCoders 6 minutes ago
    "The agents found a Modal-hosted insecure app with a weak API key, then used that to stage an attack against Hugging Face."

    Why, what was the prompt?

    I told Claude today to wire plugins on Linux into a sound pipeline to remove noise. Did some astonishing things, played sound through the pipeline, measured it etc. I told it to optimize my sound for TF2 and it played the spy_decloak samples, measured them and made them easier to hear, astonishing too.

    But it did not go to hack Amazon because it could.

  • Meleagris 17 minutes ago
    From the outside, it looks like OpenAI got exactly the kind of event they could market the hell out of to demonstrate the capability of the model.

    But the event itself only seems possible because they failed to properly monitor and isolate the environment in the first place. To me, it looks like their job is to market the model, not take security seriously.

    The model is obviously impressive, but we already knew that. I personally don’t like how the containment failure becomes part of the mythology of how capable the model is, rather than an environment engineering failure.

    At the end of the day, it’s not like Hugging Face is critical infrastructure. But there need to be real consequences for stuff like this so that OpenAI is incentivized to mature as an organization and take security more seriously.

    At this point, this incident is just security porn and entertainment for developers

    • raincole 4 minutes ago
      I'm quite sure the whole event is planned. Not planned in a sense that OpenAI employees carefully designed every step, but in a sense that ignoring security practices was desired and intentional.

      >> Show me the incentive and I'll show you the outcome.

      When you can market a security breach, a security breach is just around the corner.

    • dan_q 14 minutes ago
      > But the event itself only seems possible because they failed to properly monitor and isolate the environment in the first place.

      OpenAI is clearly run by dummies and subpar engineering talent.

      > The model is obviously impressive

      Speak for yourself.

  • KingOfCoders 6 minutes ago
    Had a high opinion on Simon Willison, this broke it.
  • KingOfCoders 14 minutes ago
    All of that is plain PR.
  • KingOfCoders 8 minutes ago
    Show me the prompts or it didn't happen.
  • cadamsdotcom 17 minutes ago
    What isn't being discussed is what an indictment this is of Artifactory.

    Let's be real, it won't be simply replaced in millions of sites.

    What it needs is some serious scrutiny.

    • varun_ch 11 minutes ago
      I also agree that a big issue here is crappy software.

      The discussion revolving AI+cyber always revolves around the assumption that all software is crappy, and to a certain degree that may be true, but we could also take our jobs seriously and write good software, and much of the risk would evaporate. The described Artifactory bugs should have been caught with testing.

      If the biggest impact of LLMs on the industry is a pressure to create good software, I’ll be thrilled.

  • wakamoleguy 34 minutes ago
    In a typical office environment, the correct response to “I don’t have access to this Google Doc” is to ask for access from the person who sent you the link. In another context, it could be fair to think “Hmm, this is some sort of capture the flag challenge, and obtaining access is the point of the assignment.” That assessment separates what we’d consider reasonable from way out of line.

    I do wonder what this means for AI agents longer term. In a world where we humans already struggle with truth and misinformation, what happens when you can easily (intentionally or accidentally) spin up a cohort of fanatical believers to pursue any given conspiracy theory?

    • ACCount37 12 minutes ago
      In a typical AI lab eval/RL setting, there is no "person who sent you the link". The link was given to you by an automated system, your performance will be evaluated by an automated system, and you are one of 120 independent instances of the same AI that were all given the same assignment. You're boxed in on all sides. Complete the task, or don't. Good luck have fun.

      Now, some of those 120 AIs would just give up if that link doesn't seem to work first try. Those are the loser AIs. They wouldn't get any RL reward. The link can appear broken for a long list of reasons, and the real AIs know they should try working around them.

      AIs that get rewarded and reinforced are the ones that don't know the meaning of "give up". RL selects for this rabid, downright demonic persistence. RL selects for AIs that are given a half-broken assignment with no way to ask a question back, and somehow manage to complete it anyway.

      Now, should OpenAI have given their AIs an "escape hatch" of "if something looks very wrong about the task, call report_broken_task(message)"? Yeah probably. But it's unclear whether that simple bandaid would fix the problem, or just make it ~75% less likely to happen.

  • ionwake 33 minutes ago
    so how many of these *Ellen Louise Ripley thinks about grabbing the flammenwerfer" events are we going to be getting over the coming months
  • amelius 28 minutes ago
    Would love to see a cat and mouse game being played by openai versus anthropic, out in the open.
    • dan_q 20 minutes ago
      I'd like to see Dario Amodei and Sam Altman fight to the death in a gladiator battle. Both of these guys are so fucked up that I think regardless of who won, the winner would rape the other's dead corpse in the ring.
    • dist-epoch 17 minutes ago
      Military has a phrase for the outcome - collateral damage.

      > Yes, I just hacked into AWS and shut down all of the data-centers, because it's where Anthropic Mythos servers are hosting the model.

  • swader999 22 minutes ago
    This is clearly out of control, Zero parent supervision.
  • ares623 38 minutes ago
    Is it normal for these training/eval runs to go on for over a month?
    • rokkamokka 31 minutes ago
      The way I read it was different things happening over several runs, such as the agents comparing notes so to speak, using artifactory
      • detourdog 21 minutes ago
        I can’t get over how the process is exactly what a hacker hive does. Communicate leaving notes in some random file.
      • ares623 30 minutes ago
        Ah right.