sh-detection-9967a81f·1 events·first seen Aliases: SH-Detection
Researchers analyze how four LLMs encode self-harm content using linear probes and contrastive direction extraction across two datasets (X-Sensitive and SH-Detection). A key finding is that self-harm information crystallizes in the final 3–7% of network layers (93–97% depth), and that probe accuracy does not correlate straightforwardly with linear separability. Gemma-3-4B shows a notably different internal representation of the contrastive self-harm direction compared to other models, with implications for detection, intervention, and LLM governance.