langevin-dynamics-101d3220·2 events·first seen Aliases: Langevin Dynamics
Researchers propose a blueprint for a thermodynamic computing stack that uses Langevin dynamics with tunable energy potentials implemented in stochastic analog hardware to run probabilistic machine learning workloads. The framework maps popular ML model classes onto hardware-native energy-based models via probabilistic graphical models, with runtime and energy analysis via theory and simulation. A preliminary experimental realization using superconducting circuits driven by thermal noise is presented. The work targets the energy and latency bottlenecks of conventional ML hardware.
This paper establishes the first global convergence theory for Wasserstein Policy Gradient (WPG), a continuous-control RL optimization method that uses optimal-transport geometry over action distributions. The authors show that the Bellman recursion structure of entropy-regularized RL induces a Polyak–Łojasiewicz (PL) geometry that substitutes for classical convexity, enabling global convergence analysis. Key technical contributions include a statewise KL representation of the soft Bellman residual, a Bellman resolvent identity linking value improvement to relative Fisher information, and a uniform log-Sobolev inequality for the evolving Gibbs policy family. The result yields geometric contraction up to discretization bias, providing theoretical grounding for WPG in continuous-action settings.