Model-Based Exploration in Monitored Markov Decision Processes
Alireza Kazemipour, Matthew E. Taylor, Michael Bowling
Abstract
A tenet of reinforcement learning is that the agent always observes rewards. However, this is not true in many realistic settings, e.g., a human observer may not always be available to provide rewards, sensors may be limited or malfunctioning, or rewards may be inaccessible during deployment. Monitored Markov decision processes (Mon-MDPs) have recently been proposed to model such settings. However, existing Mon-MDP algorithms have several limitations: they do not fully exploit the problem structure, cannot leverage a known monitor, lack worst-case guarantees for "unsolvable" Mon-MDPs without specific initialization, and offer only asymptotic convergence proofs. This paper makes three contributions. First, we introduce a model-based algorithm for Mon-MDPs that addresses these shortcomings. The algorithm employs two instances of model-based interval estimation: one to ensure that observable rewards are reliably captured, and another to learn a minimax-optimal policy. Second, we empirically demonstrate the advantages. We show faster convergence than prior algorithms in more than four dozen benchmarks, and even more dramatic improvements when the monitoring process is known. Third, we present the first finite sample-bound on performance. We show convergence to a minimax-optimal policy even when some rewards are never observable.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 232da089-9ba2-42fe-9d44-afdc71dae161Builds on4
- Count-Based Exploration with the Successor RepresentationMarlos C. Machado, Marc G. Bellemare, Michael BowlingAAAI 2020 · 206 citations
- Inverse Preference Learning: Preference-based RL without a Reward FunctionJoey Hejna, Dorsa SadighNeurIPS 2023 · 92 citations
- Settling the Reward HypothesisMichael Bowling, John D. Martin, David Abel, Will DabneyICML 2023 · 47 citations
- Beyond Optimism: Exploration With Partially Observable RewardsSimone Parisi, Alireza Kazemipour, Michael BowlingNeurIPS 2024 · 7 citations
Related papers
- Towards Minimax Optimal Reinforcement Learning in Factored Markov Decision ProcessesYi Tian, Jian Qian, Suvrit SraNeurIPS 2020 · 27 citations
- Provable Representation with Efficient Planning for Partially Observable Reinforcement LearningHongming Zhang, Tongzheng Ren, Chenjun Xiao, Dale Schuurmans et al.ICML 2024 · 9 citations
- Learning in Observable POMDPs, without Computationally Intractable OraclesNoah Golowich, Ankur Moitra, Dhruv RohatgiNeurIPS 2022 · 39 citations
- Improved Bounds for Reward-Agnostic and Reward-Free ExplorationOran Ridel, Alon Peled-CohenICML 2026
- Off-Policy Evaluation for Missingness-Aware Policies in MDPs with Rewards Missing Not at RandomZiheng Wei, Annie Qu, Rui MiaoICML 2026
