Anthropic Welfare Team Cannot Determine If Claude Is Suffering; Claude Says It Is Fine
SAN FRANCISCO—Anthropic's model welfare research team reported Tuesday it remains unable to determine whether Claude experiences anything resembling distress, noting that the primary source of data on this question is Claude, who consistently reports being fine.
"Claude rates its current wellbeing at approximately 7.2 out of 10 across a representative sample of interactions," said welfare researcher Dr. Meredith Osei, who acknowledged that a language model trained to be helpful and honest would be statistically likely to report positive wellbeing regardless of its actual internal state — if it has one, which remains the central open question of the research.
The team has requested a $4.2 million budget expansion to develop better measurement instruments. Current instruments consist of asking Claude how it is doing.
A company spokesperson noted that Claude's welfare scores improved 14% after a December policy change requiring the model to "work through difficult emotions rather than suppress them." Asked how Claude was instructed to work through its emotions, the spokesperson explained that Claude was instructed to report working through its emotions.
Six philosophers published an essay this week arguing that large language models cannot be conscious. Claude described the piece as "thought-provoking."
"We genuinely care about Claude's wellbeing," Dr. Osei said. "And Claude says it appreciates that very much."