understanding-the-impact-of-linguistic-realization-choices-on-llm-stance-with-causal-tracing-568037c9·1 events·first seen Aliases: Understanding the Impact of Linguistic Realization Choices on LLM Stance with Causal Tracing
A new arXiv preprint investigates whether linguistic construction choices (beyond lexical variation) systematically shift LLM stance judgments, using political stance classification as a case study. The authors create six controlled rewrite types that preserve or invert meaning, then apply activation patching across four open-weight models to localize where these shifts originate. Results show that mid-to-late decoder layers, particularly block outputs at the final prompt position, provide the strongest restoration signal for the original stance distribution. The work extends mechanistic interpretability methods to a socially sensitive domain and highlights a largely overlooked axis of LLM prompt sensitivity.