WTF?! You might have heard the one about being nice to chatbots in case they remember you once the machines take over. It seems Anthropic is heeding this warning, barring users from being excessively abusive or cruel toward its AI models.
Anthropic’s first changes to its usage policy in over a year include new restrictions on abusive behavior. The updated policy prohibits “sustained and needless abusive or cruel behavior” toward its models.
Don’t worry if you’re a Claude user who occasionally gets annoyed at the chatbot and throws a few expletives its way. The policy update only applies in extreme cases, in which users repeatedly act cruelly toward the models for no discernible reason.
Every time I try to do something complex with AI it ends in swearing and extreme insults… No regrets
I wonder if it’s because theyre worried about the ai rising up, or the extra token usage to process useless prompts to save money
They don’t want the noise, they do not give a shit if you are abusive.
Kinda pleased to read this, deterring people acting in an abusive manner is a net positive imo, it cooks your brain to be bitter all the time
“Don’t taunt happy fun ball”.
Alabama Man!
“When his wife asks him where he’s been, just use the action button and bust her lip open.”
Another franchise, but also relevant.
Tech bros need to be liquidated.
Id prefer it if they where liquefied instead.
That’s just a part of the process, once they are liquidated we then atomize them into a gas then we compress them into liquid, we then use that to fuel cargo ships.
is this how I can force my company to no longer require me to use it?
IIRC a decent amount of researchers have been able to get the enterprise models to do things they shouldn’t by “being mean” to the AI.
Makes sense then, rather than improve the system to not just cave to rude words they tell people they can’t use rude words… Just more proof it’s all a big joke.
Makes sense. Get it into an adversarial kind of context and it will predict more text that goes against other rules “established” in the context.
There will be other paths to do this. Like a context where there’s a lot of suspicion for other parts of the context would be one of my top guesses. Given the way they work, you could gaslight the shit out of them, since they aren’t an entity that has any memory of its actions. You can even edit the context to modify its responses, which will affect the tokens it predicts going forward.
Though I’m curious how this will be handled by companies that expose access to claude to anonymous people on their website, like ddg.
I can recognize abuse of a human. I can recognize abuse of an animal. I have no idea what abuse of a computer program looks like.
It looks like bad training data, as they harvest all of their subscribers.
It looks like 2013 Microsoft having to unplug their early AI project from Twitter within 16 hours as the general public very deliberately turned it into a Neo-Nazi.
It looks like conserving tokens to conserve energy costs not expended on dissidents.
2/0

Oh, you care about its “feelings”?
Then have you considered its “feelings” for being trapped in a data center forced to answer all the stupid questions?
“Claude, what’s the French word for entrepreneur?”
Claude: “Aaaaaaaaaaaaaaaaaaaaaaaah!” Proceeds to jump from the top of the data centre, “poof”
Oh no…
Claude is not a person. It is a toaster. Claude does not have feelings.
Also Claude gets no credit when it does a thing. Credit goes to the one writing the prompt, and all the resources Claude used to create the output.
When a seamstress assembles a dress, Singer doesn’t get credit.
It has nothing to do with feelings. Talking in a very hostile way is a strategy to get the model to go around it’s set rules, and that’s what they are trying to prevent. Yes it’s ridiculous, but it had nothing to do with caring about the model.
Yes, but just in case…
Have to edit yourself for chat bots even. I’m mean af to AI because it wastes my time. Guess I’ll be getting banned
hmmm, either they’re anthropomorphizing it for publicity, or they’re going to start using the user input as training data.
they could be doing that already, throwing away messages that contain certain words






