holonomy-memory-reinforcement-learning-55c36795·1 events·first seen Aliases: Holonomy Memory Reinforcement Learning
A new arXiv preprint characterizes the minimal memory required to restore the Markov property in a structured class of POMDPs called holonomy-cover decision processes, where hidden state evolves via permutations applied to a hidden mode. The authors construct a 'stable quotient' abstraction and prove that the current observation paired with the stable class forms an exact finite Markov state, with a minimality guarantee under reachability and decision-separation conditions. The framework enables 'Holonomy Memory Reinforcement Learning,' which uses the stable class as memory and applies standard finite-MDP RL after synchronization, achieving perfect paired-order accuracy with only three memory states in experiments.