Reinforcement Learning for Cost-Aware Markov Decision Processes
Wesley Suttle, Kaiqing Zhang, Zhuoran Yang, Ji Liu, David N. Kraemer
摘要
Ratio maximization has applications in areas as diverse as finance, reward shaping for reinforcement learning (RL), and the development of safe artificial intelligence, yet there has been very little exploration of RL algorithms for ratio maximization. This paper addresses this deficiency by introducing two new, model-free RL algorithms for solving cost-aware Markov decision processes, where the goal is to maximize the ratio of longrun average reward to long-run average cost. The first algorithm is a two-timescale scheme based on relative value iteration (RVI) Q-learning and the second is an actor-critic scheme. The paper proves almost sure convergence of the former to the globally optimal solution in the tabular case and almost sure convergence of the latter under linear function approximation for the critic. Unlike previous methods, the two algorithms provably converge for general reward and cost functions under suitable conditions. The paper also provides empirical results demonstrating promising performance and lending strong support to the theoretical results.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Finite-Time Analysis of Whittle Index based Q-Learning for Restless Multi-Armed Bandits with Neural Network Function ApproximationGuojun Xiong, Jian LiNeurIPS 2023 · 被引用 23 次
- Fractional Deep Reinforcement Learning for Age-Minimal Mobile Edge ComputingLyudong Jin, Ming Tang, Meng Zhang, Hao WangAAAI 2024 · 被引用 10 次
- Non-Asymptotic Guarantees for Average-Reward Q-Learning with Adaptive StepsizesZaiwei ChenNeurIPS 2025 · 被引用 6 次
它引用的顶会 Paper4
- Natural Policy Gradient Primal-Dual Method for Constrained Markov Decision ProcessesDongsheng Ding, Kaiqing Zhang, Tamer Basar, Mihailo R. JovanovicNeurIPS 2020 · 被引用 252 次
- A Finite-Time Analysis of Two Time-Scale Actor-Critic MethodsYue Wu, Weitong Zhang, Pan Xu, Quanquan GuNeurIPS 2020 · 被引用 189 次
- Sample Complexity of Asynchronous Q-Learning: Sharper Analysis and Variance ReductionGen Li, Yuting Wei, Yuejie Chi, Yuantao Gu 等NeurIPS 2020 · 被引用 149 次
- Leverage the Average: an Analysis of KL Regularization in Reinforcement LearningNino Vieillard, Tadashi Kozuno, Bruno Scherrer, Olivier Pietquin 等NeurIPS 2020 · 被引用 106 次
相关 Paper
- Planning and Learning in Average Risk-aware MDPsWeikai Wang, Erick DelageNeurIPS 2025 · 被引用 3 次
- Model-Free Robust Average-Reward Reinforcement LearningYue Wang, Alvaro Velasquez, George K. Atia, Ashley Prater-Bennette 等ICML 2023 · 被引用 25 次
- Two-Timescale Critic-Actor for Average Reward MDPs with Function ApproximationPrashansa Panda, Shalabh BhatnagarAAAI 2025 · 被引用 5 次
- Provably Convergent Actor-Critic in Risk-averse MARLYizhou Zhang, Eric MazumdarICML 2026
- Learning to Shape Rewards Using a Game of Two PartnersDavid Mguni, Taher Jafferjee, Jianhong Wang, Nicolas Perez Nieves 等AAAI 2023 · 被引用 17 次
