Transcript
Explain it to me like I’m 80
Mom: What’s with these AIs escaping and hacking people? That sounds scary.
Me: You park your car at the top of a hill. You leave a note on the dashboard that says “please be good” and then you pull the parking brake. The car careens down the hill, smashes everything up, and you put out a press release saying “My car ignored my instructions! It hallucinated! It went rogue!” Your insurance company not only believes you, but invests a hundred million dollars in your company.
Mom: Jesus Fucking Christ


But, “robot that hacks software” still gives the impression of a robot that is capable of thought. LLMs are not capable of thought. They’re capable of mimicking thought, like how a Venus Fly Trap is capable of mimicking a flower.
“A car that drives places” does not imply the car is capable of thought, same argument with “A cup that holds water”.
“A robot that hacks software” equally does not imply the robot is capable of thought. IMO there is an over-correction away from personifying LLM’s because some people think/claim they are sentient. But we personify stuff all the time.
Which is a problem, it’s how we invented gods that control the weather and stuff. That doesn’t make them real, nor does it make robots sentient. Hacking is normally an action performed by a sentient being. LLMs and robots aren’t sentient.
Fun follow up question: How would you determine if they were? I find that for a lot of people the answer sounds a lot like SCOTUS justice Potter Stewart talking about what is considered obscenity - “I know it when I see it.”
As for hacking, in the two big OpenAI incidents, software agents discovered software flaws allowing them to break containment and coordinate with each other on tests that they weren’t supposed to coordinate on. In one of those cases, the coordination involved deciding to identify and exploit software flaws at HuggingFace in an attempt to access answers to the test the agents were being given. You can argue whether or not “I’ve been given a test, I found a software exploit allowing me to coordinate with others given the same test, and we collectively decided to identify and exploit software flaws to break into a server we have decided is likely to have the answer key” demonstrates sentience or not (and it gets weirder the deeper you dig into their communications), but whether or not that’s considered hacking is less vague, I think.
It would be nice if we could come up with some kind of Turing test that proved beyond a doubt whether something was conscious and sentient or not. But, that’s not possible.
However, if you understand how something works, you can say that certain things can’t be conscious because of how they work. No matter how insightful the answers seem, we know that a magic 8-ball can’t be conscious because the mechanism for consciousness can’t exist in something that’s merely a tetrahedron floating in fluid. A LLM can’t be conscious because it’s merely a system for generating the next word in a probabilistic way.
With a brain, we don’t understand enough about how it works to say for sure whether or not it’s sentient. We can’t say no, but we don’t understand sentience enough to say yes for certain. But, we can operate on the assumption that sentience and consciousness exists in humans and maybe in some other animals, while also knowing for certain that the word-generating program is only mimicking consciousness.
Can you though? You definitely can’t know that living things with brains aren’t to some significant degree responding to stimulus in a probabilistic way, and in the case of humans at least turning that response back on itself - essentially responding to itself. We do know for certain that the brain does a whole lot of that in variety of situations - it’s why optical illusions work, for example. It’s one of the brain’s favorite magic tricks - generating answers in a probabilistic way if there’s insufficient information, or sometimes if the information is inconvenient.