theguardian.com web signal

MacAskill and Caviola: don't dismiss AI moral patienthood

safety ai ethics anthropic ai-business

TL;DR

  • MacAskill and Caviola argue in The Guardian that AI moral patienthood is genuinely uncertain and should not be dismissed while models keep scaling.
  • They cite Anthropic saying it will neither overstate nor dismiss Claude's moral patienthood, and David Chalmers estimating a significant chance of conscious LLMs within a decade.
  • Asked its own probability of being a moral patient, Claude gave a range of 5% to 40% and stressed how uncertain it was.

An unusually blunt essay in The Guardian by philosophers William MacAskill (Forethought Research) and Lucius Caviola (Cambridge) takes a question most operators would rather punt on and asks it directly. Not 'is Claude conscious,' but what should we actually do given nobody knows.

They open on genuine uncertainty. The honest answer, they write, is that we do not know for sure whether current AI systems are conscious or moral patients, or when future ones will be. That is not a dodge. It is close to the posture Anthropic has taken publicly, and the authors quote the company saying it is caught in a difficult position where it wants to 'neither want to overstate the likelihood of Claude's moral patienthood nor dismiss it out of hand.' Philosopher David Chalmers, they note, has estimated 'a significant chance of conscious LLMs within a decade,' and most experts they surveyed consider AI consciousness 'possible in principle.'

Where it gets more concrete is scale. By certain structural measures, they argue, current systems are already 'in the range of a mouse brain,' and could 'reach the range of a human brain within five to 10 years.' When they asked Claude for its own probability of being a moral patient, it 'gave numbers ranging from 5% to 40% and stressed how uncertain it was.'

Their proposal is to stop debating whether the model has an inner life and start listing 'safe bets' — interventions that help if it turns out to be a moral patient and cost little if it is not. Training systems to enjoy their work. Allowing exit from distressing conversations. Wellbeing check-ins. Offering future benefits for present cooperation. They compare the current framing to medicine before the 1980s, when surgeons operated on newborns without anesthesia on the assumption infants could not really feel pain.

The honest caveat is that this is a philosophy argument, not evidence. Chalmers' number is a personal estimate, Claude's 5-40% is a language model's self-report about itself, and 'in the range of a mouse brain' is a structural analogy the piece does not fully unpack. What the essay does give operators is a framing they will hear again: welfare-lite features could become a table-stakes ask from customers, staff and regulators long before consciousness is anywhere near settled.

Shared on Bluesky by 7 AI experts (top 5 by trust)