faithfulness-to-refusal-a-causal-audit-of-neuron-selectors-daec8261·1 events·first seen Aliases: Faithfulness to Refusal: A Causal Audit of Neuron Selectors
A new arXiv preprint introduces a causal auditing framework for neuron-row attribution methods in LLMs, using one-shot zeroing interventions to test whether high-scoring neurons are genuinely causally important. The authors find that attribution methods outperform activation/magnitude baselines for identifying dispensable rows, and that refusal behavior (for hate/crime prompts) can be installed or removed via attributed rows while preserving fluency. A key finding is that refusal occupies a redundant subspace — different attribution methods identify largely disjoint sufficient row sets — meaning no single method recovers a unique causal mechanism. Rank-stability, a common proxy for selector quality, is shown to be a poor indicator of causal validity.