theory-of-mind-e9212654·1 events·first seen Aliases: Theory of Mind
A new arXiv preprint finds that safety fine-tuning designed to prevent LLMs from claiming consciousness also suppresses their attribution of minds to non-human animals and natural objects, and reduces spiritual belief in model outputs. The authors show that ablating the learned safety-refusal direction or steering a 'consciousness vector' in activation space reverses these effects, recovering more human-like responses on sociological surveys covering religiosity, moral values, and well-being. Crucially, Theory of Mind capabilities remain intact, suggesting social reasoning is mechanistically independent from self-attribution of consciousness. The work raises concerns that current alignment approaches entangle targeted self-attribution suppression with culturally widespread and benign beliefs.