twin-delayed-ddpg-8d24e65a·1 events·first seen Aliases: Twin Delayed DDPG
A new arXiv preprint presents a policy distillation framework that transfers a high-performance TD3 reinforcement learning agent into a shallow decision tree surrogate for interpretable continuous control, demonstrated on the Inverted Pendulum benchmark. The method uses physics-aware features and 'Noisy Oracle Rollouts' to match teacher performance while providing global and local interpretability. A key finding is a fundamental trade-off: discretizing control induces Bang-Bang actuation and a bimodal limit cycle, though BIBO stability is preserved. The work targets safety-critical deployment contexts where regulatory compliance and human-agent trust require explainability.