Flowing Through States: Neural ODE Regularization for Reinforcement Learning
Mohamed Ghanem, Bernd Finkbeiner
Abstract
Neural networks applied to sequential decision-making tasks typically rely on latent representations of environment states. While environment dynamics dictate how semantic states evolve, the corresponding latent transitions are usually left implicit, creating a potential misalignment between the two. We propose to model latent dynamics explicitly by drawing an analogy between Markov decision process (MDP) trajectories and ordinary differential equation (ODE) flows: in both cases, the current state fully determines its successors. Building on this view, we introduce a neural ODE-based regularization method that enforces latent embeddings to follow consistent ODE flows, thereby aligning representation learning with environment dynamics. Although broadly applicable to deep learning agents, we demonstrate its effectiveness in reinforcement learning by integrating it into Actor-Critic algorithms. Our approach yields major performance gains across various standard Atari benchmarks for A2C and gridworld environments for PPO.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on4
- Data-Efficient Reinforcement Learning with Self-Predictive RepresentationsMax Schwarzer, Ankesh Anand, Rishab Goel, R. Devon Hjelm et al.ICLR 2021 · 399 citations
- TACO: Temporal Latent Action-Driven Contrastive Loss for Visual Reinforcement LearningRuijie Zheng, Xiyao Wang, Yanchao Sun, Shuang Ma et al.NeurIPS 2023 · 89 citations
- Model-based Reinforcement Learning for Semi-Markov Decision Processes with Neural ODEsJianzhun Du, Joseph Futoma, Finale Doshi-VelezNeurIPS 2020 · 63 citations
- Contrastive Behavioral Similarity Embeddings for Generalization in Reinforcement LearningRishabh Agarwal, Marlos C. Machado, Pablo Samuel Castro, Marc G. BellemareICLR 2021 · 27 citations
Related papers
- ODE-based Recurrent Model-free Reinforcement Learning for POMDPsXuanle Zhao, Duzhen Zhang, Liyuan Han, Tielin Zhang et al.NeurIPS 2023 · 18 citations
- Anamnesic Neural Differential Equations with Orthogonal Polynomial ProjectionsEdward De Brouwer, Rahul G. KrishnanICLR 2023 · 2 citations
- Learning Continuous System Dynamics from Irregularly-Sampled Partial ObservationsZijie Huang, Yizhou Sun, Wei WangNeurIPS 2020 · 103 citations
- Text Generation as Continuous Latent Dynamics via Reinforcement LearningChen JiaICML 2026
- Continuous-time Model-based Reinforcement LearningÇagatay Yildiz, Markus Heinonen, Harri LähdesmäkiICML 2021 · 72 citations
