analysing-self-harm-representations-in-language-models-a-cross-architecture-study-25428cd2·1 events·first seen Aliases: Analysing Self-Harm Representations in Language Models: a Cross-Architecture Study
Researchers analyze how four LLMs encode self-harm content using linear probes and contrastive direction extraction across two datasets (X-Sensitive and SH-Detection). A key finding is that self-harm information crystallizes in the final 3–7% of network layers (93–97% depth), and that probe accuracy does not correlate straightforwardly with linear separability. Gemma-3-4B shows a notably different internal representation of the contrastive self-harm direction compared to other models, with implications for detection, intervention, and LLM governance.