Autoregressive Policy Optimization for Constrained Allocation Tasks
David Winkel, Niklas Strauß, Maximilian Bernhard, Zongyue Li, Thomas Seidl, Matthias Schubert
Abstract
Allocation tasks represent a class of problems where a limited amount of resources must be allocated to a set of entities at each time step. Prominent examples of this task include portfolio optimization or distributing computational workloads across servers. Allocation tasks are typically bound by linear constraints describing practical requirements that have to be strictly fulfilled at all times. In portfolio optimization, for example, investors may be obligated to allocate less than 30% of the funds into a certain industrial sector in any investment period. Such constraints restrict the action space of allowed allocations in intricate ways, which makes learning a policy that avoids constraint violations difficult. In this paper, we propose a new method for constrained allocation tasks based on an autoregressive process to sequentially sample allocations for each entity. In addition, we introduce a novel de-biasing mechanism to counter the initial bias caused by sequential sampling. We demonstrate the superior performance of our approach compared to a variety of Constrained Reinforcement Learning (CRL) methods on three distinct constrained allocation tasks: portfolio optimization, computational workload distribution, and a synthetic allocation benchmark. Our code is available at: https://github.com/niklasdbs/paspo
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 36b7d009-dd45-4cd9-aa91-288349150788Builds on6
- Responsive Safety in Reinforcement Learning by PID Lagrangian MethodsAdam Stooke, Joshua Achiam, Pieter AbbeelICML 2020 · 403 citations
- First Order Constrained Optimization in Policy SpaceYiming Zhang, Quan Vuong, Keith W. RossNeurIPS 2020 · 238 citations
- IPO: Interior-Point Policy Optimization under ConstraintsYongshuai Liu, Jiaxin Ding, Xin LiuAAAI 2020 · 231 citations
- Constrained Update Projection Approach to Safe Policy OptimizationLong Yang, Jiaming Ji, Juntao Dai, Linrui Zhang et al.NeurIPS 2022 · 95 citations
- Solving Online Threat Screening Games using Constrained Action Space Reinforcement LearningSanket Shah, Arunesh Sinha, Pradeep Varakantham, Andrew Perrault et al.AAAI 2020 · 14 citations
Related papers
- Efficient Action-Constrained Reinforcement Learning via Acceptance-Rejection Method and Augmented MDPsWei Hung, Shao-Hua Sun, Ping-Chun HsiehICLR 2025
- Multi Agent Reinforcement Learning for Sequential Satellite Assignment ProblemsJoshua Holder, Natasha Jaques, Mehran MesbahiAAAI 2025 · 5 citations
- Gradient-Adaptive Pareto Optimization for Constrained Reinforcement LearningZixian Zhou, Mengda Huang, Feiyang Pan, Jia He et al.AAAI 2023 · 11 citations
- Quantile Constrained Reinforcement Learning: A Reinforcement Learning Framework Constraining Outage ProbabilityWhiyoung Jung, Myungsik Cho, Jongeui Park, Youngchul SungNeurIPS 2022 · 12 citations
- Extreme Value Policy Optimization for Safe Reinforcement LearningShiqing Gao, Yihang Zhou, Shuai Shao, Haoyu Luo et al.ICML 2025
