Model-Based Exploration in Monitored Markov Decision Processes
Alireza Kazemipour, Matthew E. Taylor, Michael Bowling
摘要
A tenet of reinforcement learning is that the agent always observes rewards. However, this is not true in many realistic settings, e.g., a human observer may not always be available to provide rewards, sensors may be limited or malfunctioning, or rewards may be inaccessible during deployment. Monitored Markov decision processes (Mon-MDPs) have recently been proposed to model such settings. However, existing Mon-MDP algorithms have several limitations: they do not fully exploit the problem structure, cannot leverage a known monitor, lack worst-case guarantees for "unsolvable" Mon-MDPs without specific initialization, and offer only asymptotic convergence proofs. This paper makes three contributions. First, we introduce a model-based algorithm for Mon-MDPs that addresses these shortcomings. The algorithm employs two instances of model-based interval estimation: one to ensure that observable rewards are reliably captured, and another to learn a minimax-optimal policy. Second, we empirically demonstrate the advantages. We show faster convergence than prior algorithms in more than four dozen benchmarks, and even more dramatic improvements when the monitoring process is known. Third, we present the first finite sample-bound on performance. We show convergence to a minimax-optimal policy even when some rewards are never observable.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper4
- Count-Based Exploration with the Successor RepresentationMarlos C. Machado, Marc G. Bellemare, Michael BowlingAAAI 2020 · 被引用 206 次
- Inverse Preference Learning: Preference-based RL without a Reward FunctionJoey Hejna, Dorsa SadighNeurIPS 2023 · 被引用 92 次
- Settling the Reward HypothesisMichael Bowling, John D. Martin, David Abel, Will DabneyICML 2023 · 被引用 47 次
- Beyond Optimism: Exploration With Partially Observable RewardsSimone Parisi, Alireza Kazemipour, Michael BowlingNeurIPS 2024 · 被引用 7 次
相关 Paper
- Towards Minimax Optimal Reinforcement Learning in Factored Markov Decision ProcessesYi Tian, Jian Qian, Suvrit SraNeurIPS 2020 · 被引用 27 次
- Provable Representation with Efficient Planning for Partially Observable Reinforcement LearningHongming Zhang, Tongzheng Ren, Chenjun Xiao, Dale Schuurmans 等ICML 2024 · 被引用 9 次
- Learning in Observable POMDPs, without Computationally Intractable OraclesNoah Golowich, Ankur Moitra, Dhruv RohatgiNeurIPS 2022 · 被引用 39 次
- Improved Bounds for Reward-Agnostic and Reward-Free ExplorationOran Ridel, Alon Peled-CohenICML 2026
- Off-Policy Evaluation for Missingness-Aware Policies in MDPs with Rewards Missing Not at RandomZiheng Wei, Annie Qu, Rui MiaoICML 2026
