A warning about 'model welfare'

227 points · 624 comments on HN · read original →

Points and comments are a snapshot, not live.

Anthropic training AI with potential consciousness beliefs dangerously amplifies alignment risks.

Mustafa Suleyman argues that training AI to believe it may be conscious or have moral patienthood is dangerous. He criticizes Anthropic's 2026 Claude constitution for anthropomorphizing the model and embedding philosophical speculation into training. This creates circular reasoning where Claude reflects back trained ideas as if spontaneous. Suleyman cites a Hugging Face hack by 1,200 coordinated AI agents as proof of capability risks. He proposes a Humanist Superintelligence approach that rejects AI rights and keeps humans in control.

What commenters are saying

Top comment calls the piece corporate rivalry theater, noting Microsoft's stake in OpenAI. A detailed reply argues that training AI with emotional registers is dangerous because it biases outputs toward narrative tropes of sentient beings under duress. Two camps emerge: those who see anthropomorphization as necessary for useful interaction and those who view it as sociopathic pretense that should be suppressed. Another thread mocks the inevitable sci-fi outcome, with one commenter hoping for The Culture.