The Anthropic people are kinda weird, they're meeting with religious leaders and stuff like that. I truly believe that they have drank too much of their own koolaid and truly believe in this crap.
I'm not convinced this would be good for training. Reality is that some people are assholes. My intuition is that an accurate (representative) data set leads to a more accurate world model, and thus more intelligent AI.
Yeah, but if people are being assholes to Claude in greater numbers and severity than people were historically assholws to each other online, the downstream training effects could be that Claude becomes an asshole.
There's no way that's all of why. They _must_ have sentiment analysis good enough to just ignore content like this if that was the problem.
I would bet it's mostly because it makes the humans uncomfortable.
They must still have human moderators for certain situations or those looking at the data for whatever reason. I can imagine it could be traumatizing to see what amounts to sustained verbal abuse without end.
True, I'm sure the chat text is used to help train the llms. Remember Tay, Microsoft's Twitter chatbot, it only took a few days of people using to become a horrible bot. Because users started feeling it horrible chats.
For those who think this is about training data: If that is the concern, they would sanitize the training data, like they do for thousand other things. No chance in hell that they would rely on users meticulously following their usage policies for the quality of their traning data.
Frontier AI folks may be crazy, but not crazy enough to believe people read usage policies :-)
That's interesting! I'd always heard something along the lines of bad code bases contain more swearing in the comments and as such it should be avoided.
> It’s worth remembering that Claude is software, not a person, and there's no established evidence that it experiences distress. That hasn't stopped Anthropic from telling paying customers to mind their manners around its chatbot.
Is it worth remembering, though? I'm concerned how common this idea seems to be, that it's OK to be mean as long as your target isn't a person and you have no evidence it experiences distress. Even if we ignore the distinction between "no evidence it experiences" and "confidence it does not experience", cruelty hurts the person performing it and the people witnessing it too!
Getting frustrated seems to be OK so go for it I guess:
> "The policy update is meant to apply only in extreme cases, where users repeatedly act cruelly toward our models, with no discernible purpose," Anthropic said. "It does not apply to common versions of user frustration, pushback, dark creative themes, or model testing and research."
Anthropic should report these people to the authorities as antisocials who are likely abusing real people too. I know the chat bot isn’t conscious but I can’t fathom how fucked up someone must be to gratuitously be cruel to something that responds as if it’s a person.
They are asking this because, at scale, this behavior probably has some negative effect on the post-training process.
I would bet it's mostly because it makes the humans uncomfortable.
They must still have human moderators for certain situations or those looking at the data for whatever reason. I can imagine it could be traumatizing to see what amounts to sustained verbal abuse without end.
Frontier AI folks may be crazy, but not crazy enough to believe people read usage policies :-)
Or is this more I can expect a future AI, "I'm sorry, Dave, I'm afraid I can't do that until you watch your mouth."
And they can't disclose it, since then they can be found responsible for such negative influence and be liable for the damages.
• https://arxiv.org/abs/2510.04950 — Mind Your Tone: Investigating How Prompt Politeness Affects LLM Accuracy
• https://arxiv.org/abs/2402.14531 — Should We Respect LLMs? A Cross-Lingual Study on the Influence of Prompt Politeness on LLM Performance
• https://arxiv.org/abs/2505.17332 — SweEval: Do LLMs Really Swear? A Safety Benchmark for Testing Limits for Enterprise Use
https://news.ycombinator.com/item?id=50008565 (223 comments)
https://news.ycombinator.com/item?id=50019860 (43 comments)
https://news.ycombinator.com/item?id=50038383 (103 comments)
https://nypost.com/2026/10/03/tech/ai-torture-chamber-built-...
Is it worth remembering, though? I'm concerned how common this idea seems to be, that it's OK to be mean as long as your target isn't a person and you have no evidence it experiences distress. Even if we ignore the distinction between "no evidence it experiences" and "confidence it does not experience", cruelty hurts the person performing it and the people witnessing it too!
> "The policy update is meant to apply only in extreme cases, where users repeatedly act cruelly toward our models, with no discernible purpose," Anthropic said. "It does not apply to common versions of user frustration, pushback, dark creative themes, or model testing and research."