knowledge-and-gradient-guided-reinforcement-learning-for-parametrized-action-markov-decision-processes-8447d20e·1 events·first seen Aliases: Knowledge- and Gradient-Guided Reinforcement Learning for Parametrized Action Markov Decision Processes
A new arXiv preprint introduces KGRL (Knowledge- and Gradient-Guided Reinforcement Learning), a neuro-symbolic algorithm for Parametrized Action Markov Decision Processes (PAMDPs) where each decision involves both a symbolic action and continuous numerical parameters. KGRL uses a Datalog knowledge base to prune infeasible actions and constrain parameter spaces, then applies gradient-based refinement to estimate optimal parameters during training and deployment. The approach improves sample efficiency and episodic return over state-of-the-art PAMDP baselines, and provides local procedural explanations for its decisions.