what-do-reward-models-memorize--69d6a5b7·1 events·first seen Aliases: What do Reward Models Memorize?
A new arXiv paper measures counterfactual memorization in discriminatively trained reward models (RMs) across two human preference datasets. The authors find that RMs misallocate memorization to easy high-margin pairs, learn dataset-specific shortcuts (e.g., model identity, user sampling strategy), and overgeneralize simple heuristics like response length and compliance. The findings suggest current discriminative RM training produces biased models that fail to judge response quality in context-dependent scenarios, with direct implications for RLHF pipelines.