Generative Modelling of Stochastic Actions with Arbitrary Constraints in Reinforcement Learning
Changyu Chen, Ramesha Karunasena, Thanh Hong Nguyen, Arunesh Sinha, Pradeep Varakantham
Abstract
Many problems in Reinforcement Learning (RL) seek an optimal policy with large discrete multidimensional yet unordered action spaces; these include problems in randomized allocation of resources such as placements of multiple security resources and emergency response units, etc. A challenge in this setting is that the underlying action space is categorical (discrete and unordered) and large, for which existing RL methods do not perform well. Moreover, these problems require validity of the realized action (allocation); this validity constraint is often difficult to express compactly in a closed mathematical form. The allocation nature of the problem also prefers stochastic optimal policies, if one exists. In this work, we address these challenges by (1) applying a (state) conditional normalizing flow to compactly represent the stochastic policy -the compactness arises due to the network only producing one sampled action and the corresponding log probability of the action, which is then used by an actor-critic method; and (2) employing an invalid action rejection method (via a valid action oracle) to update the base policy. The action rejection is enabled by a modified policy gradient that we derive. Finally, we conduct extensive experiments to show the scalability of our approach compared to prior methods and the ability to enforce arbitrary state-conditional constraints on the support of the distribution of actions in any state 1 .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b10c0ce7-32aa-4faa-88cf-4d88978a1608Cited by top-tier papers6
- Reward Penalties on Augmented States for Solving Richly Constrained RL EffectivelyHao Jiang, Tien Mai, Pradeep Varakantham, Huy HoangAAAI 2024 · 2 citations
- Leveraging Constraint Violation Signals for Action Constrained Reinforcement LearningJanaka Chathuranga Brahmanage, Jiajing Ling, Akshat KumarAAAI 2025 · 2 citations
- Improving and Accelerating Offline RL in Large Discrete Action Spaces with Structured Policy InitializationMatthew Landers, Taylor W. Killian, Tom Hartvigsen, Afsaneh DoryabICLR 2026 · 1 citation
- APC-RL: Exceeding data-driven behavior priors with adaptive policy compositionFinn Rietz, Pedro Zuidberg Dos Martires, Johannes A. StorkICLR 2026
- Action-Constrained Imitation LearningChia-Han Yeh, Tse-Sheng Nan, Risto Vuorio, Wei Hung et al.ICML 2025
Builds on6
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- Score-Based Generative Modeling through Stochastic Differential EquationsYang Song, Jascha Sohl-Dickstein, Diederik P. Kingma, Abhishek Kumar et al.ICLR 2021 · 1,270 citations
- PettingZoo: Gym for Multi-Agent Reinforcement LearningJ. K. Terry, Benjamin Black, Nathaniel Grammel, Mario Jayakumar et al.NeurIPS 2021 · 478 citations
- Discretizing Continuous Action Space for On-Policy OptimizationYunhao Tang, Shipra AgrawalAAAI 2020 · 150 citations
- DC3: A learning method for optimization with hard constraintsPriya L. Donti, David Rolnick, J. Zico KolterICLR 2021 · 64 citations
Related papers
- FlowPG: Action-constrained Policy Gradient with Normalizing FlowsJanaka Chathuranga Brahmanage, Jiajing Ling, Akshat KumarNeurIPS 2023 · 16 citations
- Maximum Entropy Reinforcement Learning via Energy-Based Normalizing FlowChen-Hao Chao, Chien Feng, Wei-Fang Sun, Cheng-Kuang Lee et al.NeurIPS 2024 · 29 citations
- Latent Spherical Flow Policy for Reinforcement Learning with Combinatorial ActionsLingkai Kong, Anagha Satish, Hezi Jiang, Akseli Kangaslahti et al.ICML 2026 · 1 citation
- Dynamic Neighborhood Construction for Structured Large Discrete Action SpacesFabian Akkerman, Julius Luy, Wouter van Heeswijk, Maximilian SchifferICLR 2024 · 5 citations
- Solving Continuous Mean Field Games: Deep Reinforcement Learning for Non-Stationary DynamicsLorenzo Magnino, Kai Shao, Zida Wu, Jiacheng Shen et al.NeurIPS 2025 · 4 citations
