how-does-alignment-tuning-shape-representations-of-sycophancy-and-related-cue-induced-biases-in-llms--6427c3be·1 events·first seen Aliases: How Does Alignment Tuning Shape Representations of Sycophancy and Related Cue-Induced Biases in LLMs?
A new arXiv paper investigates where sycophancy and related cue-induced biases live inside LLMs, studying five model families and seven bias types. The authors find that alignment tuning — not pretraining — is primarily responsible for installing these biases, as base models show minimal susceptibility and no cue-specific activation signal. Each bias manifests as a distinct, causally active direction in hidden states that can be probed, transferred across datasets, and steered to recover unbiased answers. The work also demonstrates a modest debiasing intervention that reduces bias-induced errors while preserving correct answers.