FlowPG: Action-constrained Policy Gradient with Normalizing Flows
Janaka Chathuranga Brahmanage, Jiajing Ling, Akshat Kumar
摘要
Action-constrained reinforcement learning (ACRL) is a popular approach for solving safety-critical and resource-allocation related decision making problems. A major challenge in ACRL is to ensure agent taking a valid action satisfying constraints in each RL step. Commonly used approach of using a projection layer on top of the policy network requires solving an optimization program which can result in longer training time, slow convergence, and zero gradient problem. To address this, first we use a normalizing flow model to learn an invertible, differentiable mapping between the feasible action space and the support of a simple distribution on a latent variable, such as Gaussian. Second, learning the flow model requires sampling from the feasible action space, which is also challenging. We develop multiple methods, based on Hamiltonian Monte-Carlo and probabilistic sentential decision diagrams for such action sampling for convex and non-convex constraints. Third, we integrate the learned normalizing flow with the DDPG algorithm. By design, a well-trained normalizing flow will transform policy output into a valid action without requiring an optimization solver. Empirically, our approach results in significantly fewer constraint violations (upto an order-of-magnitude for several instances) and is multiple times faster on a variety of continuous control tasks. 37th Conference on Neural Information Processing Systems (NeurIPS 2023).
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper8
- GenPO: Generative Diffusion Models Meet On-Policy Reinforcement LearningShutong Ding, Ke Hu, Shan Zhong, Haoyang Luo 等NeurIPS 2025 · 被引用 22 次
- Leveraging Constraint Violation Signals for Action Constrained Reinforcement LearningJanaka Chathuranga Brahmanage, Jiajing Ling, Akshat KumarAAAI 2025 · 被引用 2 次
- Cross-Domain Policy Optimization via Bellman Consistency and Hybrid CriticsMing-Hong Chen, Kuan-Chen Pan, You-De Huang, Xi Liu 等ICLR 2026 · 被引用 1 次
- Improving Stochastic Action-Constrained Reinforcement Learning via Truncated DistributionsRoland Stolz, Michael Eichelbeck, Matthias AlthoffAAAI 2026 · 被引用 1 次
- Offline Reinforcement Learning with Generative Trajectory PoliciesXinsong Feng, Leshu Tang, Chenan Wang, Haipeng ChenICML 2026 · 被引用 1 次
相关 Paper
- Efficient Action-Constrained Reinforcement Learning via Acceptance-Rejection Method and Augmented MDPsWei Hung, Shao-Hua Sun, Ping-Chun HsiehICLR 2025
- Generative Modelling of Stochastic Actions with Arbitrary Constraints in Reinforcement LearningChangyu Chen, Ramesha Karunasena, Thanh Hong Nguyen, Arunesh Sinha 等NeurIPS 2023 · 被引用 16 次
- PolyFlow: Safe and Efficient Polytope-Constrained Flow Matching with Constraint Embedding and Projection-free UpdateJianming Ma, Qiyue Yang, Yang Zhang, Liyun Yan 等ICML 2026 · 被引用 1 次
- Differentiable Causal Discovery from Interventional DataPhilippe Brouillard, Sébastien Lachapelle, Alexandre Lacoste, Simon Lacoste-Julien 等NeurIPS 2020 · 被引用 295 次
- Constrained Variational Policy Optimization for Safe Reinforcement LearningZuxin Liu, Zhepeng Cen, Vladislav Isenbaev, Wei Liu 等ICML 2022 · 被引用 112 次
