enough-is-as-good-as-a-feast-a-comprehensive-analysis-of-how-reinforcement-learning-mitigates-task-conflicts-in-llms-df546605·1 events·first seen Aliases: Enough is as good as a feast: A Comprehensive Analysis of How Reinforcement Learning Mitigates Task Conflicts in LLMs
A new arXiv preprint systematically compares model merging behavior between RL-trained and SFT-trained LLMs across five tasks, finding that RL significantly reduces task conflicts and performance degradation post-merge. The authors identify three mechanisms: smaller gradient updates from on-policy training, an optimization objective that limits conflict parameter updates at convergence, and joint positive/negative example optimization that steers models toward unbiased task-specific parameter subspaces. The findings suggest RL training paradigms are inherently better suited for model merging workflows, with implications for how practitioners choose training methods when multi-task consolidation is a goal.