active-offline-to-online-reinforcement-learning-ef39c1a7·1 events·first seen Aliases: Active Offline-to-Online Reinforcement Learning
A new arXiv preprint introduces an active policy selection framework for offline-to-online reinforcement learning (O2O-RL), addressing the problem of choosing which candidate policy to fine-tune when online interactions are scarce or risky. The method uses upper-confidence bounds derived from locally linear performance forecasts to balance exploration (policy evaluation) against exploitation (fine-tuning). Experiments show consistent outperformance over existing O2O-RL baselines. The authors claim this is the first work to formally address active policy selection in the O2O-RL setting.