dadiff-59c8934c·1 events·first seen Aliases: DADiff
DADiff is a new diffusion-based framework for online dynamics adaptation in reinforcement learning, addressing the domain transfer problem by leveraging discrepancies between source and target domain generative trajectories during next-state prediction. The authors develop both reward modification and data selection variants, and provide theoretical bounds on policy performance differences in terms of generative trajectory deviation. Experiments across environments with various domain shifts show superior performance over existing approaches. Code is publicly released.