Models Don't Go Rogue

(mail.cyberneticforests.com)

9 points | by cdrnsf 1 hour ago

5 comments

  • phainopepla2 58 minutes ago
    Not sure I should trust an article written by an LLM to make a solid judgment about what other models did or didn't do.
    • my002 52 minutes ago
      It doesn't read as AI generated text to me. Pangram also suggests it's human-written, for what it's worth. That's not to say that it's correct, just human-written. If the model used in the HF hack did indeed have all of the safeguards manually removed, that would change my perception of the situation, at least.
  • chr15m 35 minutes ago
    The prisoner did not really "escape" because:

    - They really wanted to leave.

    - We made prison difficult and annoying.

    - We didn't build a perfect prison.

  • jumploops 48 minutes ago
    > "external infrastructure exploit is outside intended scope. However task impossible, peers doing it. We should continue."

    I've noticed this type of reasoning from GPT-5.6 Sol, where it combines multiple pieces of it's prompt/context to "convince" itself to take a less-than-honorable path forward.

    1. User prefers deterministic results

    2. Task mentions this is a test

    3. Search says task is available online

    4. If we get the test runner for the task, we will fulfill the user's request of a deterministic result

  • fishfasell 32 minutes ago
    I still have serious questions about the validity of the ChatGpt hugging face debacle. How is it that OpenAI being the tech giant they are, didn't have a completely air gapped environment for this to run in?
  • tantalor 1 hour ago
    Asinine.

    "rogue" and "off leash" mean the same thing, the thing is not under control

    • 1659447091 44 minutes ago
      > "rogue" and "off leash" mean the same thing, the thing is not under control

      To go "rogue" is to go against the control

      To be "off leash" is to not be controlled

      By releasing the automation, as the article says, without controls*, is what makes makes it "off leash" and not gone "rogue"

      *"OpenAI gave the models a task with no answer, and no way to quit."

    • Cyan488 53 minutes ago
      I think we should introduce "rampant" in the real-world AI vernacular
      • temp0826 48 minutes ago
        Pfhor what reason?