Timing is Everything: Learning to Act Selectively with Costly Actions and Budgetary Constraints
David Henry Mguni, Aivar Sootla, Juliusz Ziomek, Oliver Slumbers, Zipeng Dai, Kun Shao, Jun Wang
摘要
Many real-world settings involve costs for performing actions; transaction costs in financial systems and fuel costs being common examples. In these settings, performing actions at each time step quickly accumulates costs leading to vastly suboptimal outcomes. Additionally, repeatedly acting produces wear and tear and ultimately, damage. Determining when to act is crucial for achieving successful outcomes and yet, the challenge of efficiently learning to behave optimally when actions incur minimally bounded costs remains unresolved. In this paper, we introduce a reinforcement learning (RL) framework named Learnable Impulse Control Reinforcement Algorithm (LICRA), for learning to optimally select both when to act and which actions to take when actions incur costs. At the core of LICRA is a nested structure that combines RL and a form of policy known as impulse control which learns to maximise objectives when actions incur costs. We prove that LICRA, which seamlessly adopts any RL method, converges to policies that optimally select when to perform actions and their optimal magnitudes. We then augment LICRA to handle problems in which the agent can perform at most k<\infty actions and more generally, faces a budget constraint. We show LICRA learns the optimal value function and ensures budget constraints are satisfied almost surely. We demonstrate empirically LICRA's superior performance against benchmark RL methods in OpenAI gym's Lunar Lander and in Highway environments and a variant of the Merton portfolio problem within finance.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Learning to Shape Rewards Using a Game of Two PartnersDavid Mguni, Taher Jafferjee, Jianhong Wang, Nicolas Perez Nieves 等AAAI 2023 · 被引用 17 次
- MANSA: Learning Fast and Slow in Multi-Agent SystemsDavid Henry Mguni, Haojun Chen, Taher Jafferjee, Jianhong Wang 等ICML 2023 · 被引用 4 次
- Learning Robust Multi-Agent Policies via Selective Adversarial Fault InductionDavid H Mguni, Yaqi Sun, Haojun Chen, Wanrong Yang 等ICML 2026
它引用的顶会 Paper1
相关 Paper
- COptiDICE: Offline Constrained Reinforcement Learning via Stationary Distribution Correction EstimationJongmin Lee, Cosmin Paduraru, Daniel J. Mankowitz, Nicolas Heess 等ICLR 2022 · 被引用 84 次
- Gradient-Adaptive Pareto Optimization for Constrained Reinforcement LearningZixian Zhou, Mengda Huang, Feiyang Pan, Jia He 等AAAI 2023 · 被引用 11 次
- Efficient Action-Constrained Reinforcement Learning via Acceptance-Rejection Method and Augmented MDPsWei Hung, Shao-Hua Sun, Ping-Chun HsiehICLR 2025
- Dynamic allocation of limited memory resources in reinforcement learningNisheet Patel, Luigi Acerbi, Alexandre PougetNeurIPS 2020 · 被引用 6 次
- Active Reinforcement Learning Strategies for Offline Policy ImprovementAmbedkar Dukkipati, Ranga Shaarad Ayyagari, Bodhisattwa Dasgupta, Parag Dutta 等AAAI 2025
