- cross-posted to:
- [email protected]
- cross-posted to:
- [email protected]
Once you understand that these are chatbots that were designed to complete challenges like this, using tactics like this, you can understand that the chatbots didn’t “go rogue.” They did what they were designed to do, and because OpenAI ran them with inadequate supervision (without a “human in the loop” that checked each iteration through the Python loop to ensure it hadn’t gone off the rails), they trashed a competitor’s servers.
Designing autonomous, malicious software is generally considered irresponsible and dangerous. If you showed up at Defcon and gave a talk about how your autonomous malware did something unexpected and damaged someone else’s computers, the first question from the audience would be “Why are you so shit at making secure sandboxes?” It wouldn’t be “How are you so awesome at making hacking tools?”
The fact that OpenAI is making it much easier for unskilled people to break into and damage servers is indeed very bad news, but it’s not new bad news. Irresponsible parties have been doing this for years, most notably the NSA…
…
Riley had a very good way of summarizing this: “LLMs are real, AI is fake.” LLMs – chatbots trained on things like CTF logs that can break into servers – are real. They’re on a continuum with other hacking tools that have been steadily demonstrating the fragility of the modern digital world, albeit without inspiring anyone in power to do anything about it.
“AI” – chatbots that wake up, “set their own goals,” and “spontaneously” start hacking servers – is fake. It doesn’t have “a 10% chance of ending the human race.” The Hugging Face hack isn’t a mysterious, supernatural occurrence. It’s a Python loop and a chatbot. The people responsible didn’t accidentally create god: they created autonomous malicious software and then failed to closely monitor it, resulting in it doing something both foreseeable and bad.
It’s fine to worry about this new suite of tools that give even stupider people the ability to trash even more computers. You should worry about that – and demand better security practices from firms and governments, including a blanket prohibition on NOBUS-style vulnerability hoarding. That’s a productive kind of worrying, with a chance of addressing your area of concern. It’s infinitely more reasonable than locking yourself in the toilet with a flashlight and saying “Ayyyyy Eyyyyyye” into the mirror until you wet yourself.


They did not intend for it to hack those websites, it realized that was a way of achieving its goals even though it’s specifically designed to be ethical. That can easily be called going rogue and I see no issue with it. They did not design it to pass the benchmark by hacking the website the benchmark was on. They hardly really design llm’s.
As the article said you can replace rogue agents with unpredictable computer programs if you really want, I don’t see the necessity.
Regardless that’s a major goalpost shift.
I don’t presume Anthropic is designing anything to be ethical. First, the chatbot is unpredictable by design. Randomness is built in. Second, if they didn’t intend for bad things to happen, why didn’t an employee babysit it?
Speaking of goalpost shifts: “Rogue agent” is the clickbait title across multiple articles. Remember you said nobody was misrepresenting the AIs. The still-inaccurate clarification (the unsupervised chatbot functioned exactly as it was designed to function, and the script running it allowed it to execute exactly the commands Anthropic wanted) doesn’t help the article the clarification is from. It certainly doesn’t absolve every other author.
Then why do they put any safeguards of any sort in? They put a lot of work into this, also this is openai.
Randomness is built in but this behavior was not designed, it was unpredictable and random, which is, yes, unpredictable. The goal isn’t to make it unpredictable and temperature controls exist for a reason.
They did they just weren’t paying enough attention, this was a benchmark. You have to have someone check the results to be useful.
I reject that rogue is a goalpost shift or inaccurate. I think calling this rogue behavior is accurate.
rogue /rōg/ noun
It did act that way, no? I did not shift goalposts. Remember the quote was