FlowPG: Action-constrained Policy Gradient with Normalizing Flows
Janaka Chathuranga Brahmanage, Jiajing Ling, Akshat Kumar
Abstract
Action-constrained reinforcement learning (ACRL) is a popular approach for solving safety-critical and resource-allocation related decision making problems. A major challenge in ACRL is to ensure agent taking a valid action satisfying constraints in each RL step. Commonly used approach of using a projection layer on top of the policy network requires solving an optimization program which can result in longer training time, slow convergence, and zero gradient problem. To address this, first we use a normalizing flow model to learn an invertible, differentiable mapping between the feasible action space and the support of a simple distribution on a latent variable, such as Gaussian. Second, learning the flow model requires sampling from the feasible action space, which is also challenging. We develop multiple methods, based on Hamiltonian Monte-Carlo and probabilistic sentential decision diagrams for such action sampling for convex and non-convex constraints. Third, we integrate the learned normalizing flow with the DDPG algorithm. By design, a well-trained normalizing flow will transform policy output into a valid action without requiring an optimization solver. Empirically, our approach results in significantly fewer constraint violations (upto an order-of-magnitude for several instances) and is multiple times faster on a variety of continuous control tasks. 37th Conference on Neural Information Processing Systems (NeurIPS 2023).
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 9de1a17a-6961-4a35-9294-df0e07866bb5Cited by top-tier papers8
- GenPO: Generative Diffusion Models Meet On-Policy Reinforcement LearningShutong Ding, Ke Hu, Shan Zhong, Haoyang Luo et al.NeurIPS 2025 · 22 citations
- Leveraging Constraint Violation Signals for Action Constrained Reinforcement LearningJanaka Chathuranga Brahmanage, Jiajing Ling, Akshat KumarAAAI 2025 · 2 citations
- Cross-Domain Policy Optimization via Bellman Consistency and Hybrid CriticsMing-Hong Chen, Kuan-Chen Pan, You-De Huang, Xi Liu et al.ICLR 2026 · 1 citation
- Improving Stochastic Action-Constrained Reinforcement Learning via Truncated DistributionsRoland Stolz, Michael Eichelbeck, Matthias AlthoffAAAI 2026 · 1 citation
- Offline Reinforcement Learning with Generative Trajectory PoliciesXinsong Feng, Leshu Tang, Chenan Wang, Haipeng ChenICML 2026 · 1 citation
Related papers
- Efficient Action-Constrained Reinforcement Learning via Acceptance-Rejection Method and Augmented MDPsWei Hung, Shao-Hua Sun, Ping-Chun HsiehICLR 2025
- Generative Modelling of Stochastic Actions with Arbitrary Constraints in Reinforcement LearningChangyu Chen, Ramesha Karunasena, Thanh Hong Nguyen, Arunesh Sinha et al.NeurIPS 2023 · 16 citations
- PolyFlow: Safe and Efficient Polytope-Constrained Flow Matching with Constraint Embedding and Projection-free UpdateJianming Ma, Qiyue Yang, Yang Zhang, Liyun Yan et al.ICML 2026 · 1 citation
- Differentiable Causal Discovery from Interventional DataPhilippe Brouillard, Sébastien Lachapelle, Alexandre Lacoste, Simon Lacoste-Julien et al.NeurIPS 2020 · 295 citations
- Constrained Variational Policy Optimization for Safe Reinforcement LearningZuxin Liu, Zhepeng Cen, Vladislav Isenbaev, Wei Liu et al.ICML 2022 · 112 citations
