resist-and-update-counterfactual-report-coordinates-for-incentive-compatible-llms-ddf404b1·1 events·first seen Aliases: Resist and Update: Counterfactual Report Coordinates for Incentive-Compatible LLMs
A new arXiv paper introduces Counterfactual Report Coordinates (CRC), a method for identifying and controlling low-rank activation-space coordinates governing a model's answer, confidence, and caveat outputs, with the goal of making LLMs internally incentive-compatible — resistant to sycophantic pressure while remaining responsive to genuine evidence. The authors use causal interchange interventions (not probe accuracy) to identify these coordinates and introduce a training-free two-pass clamp that achieves near-perfect resist and update scores on a Bayesian-witness benchmark. Results transfer across three model families and to SycophancyEval, though the deployable single-pass compilation is lossy. The work frames activation-level counterfactual incentive-invariance as a structural primitive for alignment research.