Learning to Explore in POMDPs with Informational Rewards
Annie Xie, Logan M. Bhamidipaty, Evan Zheran Liu, Joey Hong, Sergey Levine, Chelsea Finn
Abstract
Standard exploration methods typically rely on random coverage of the state space or coveragepromoting exploration bonuses. However, in partially observed settings, the biggest exploration challenge is often posed by the need to discover information-gathering strategies-e.g., an agent that has to navigate to a location in traffic might learn to first check traffic conditions and then choose a route. In this work, we design a POMDP agent that gathers information about the hidden state, using ideas from the meta-exploration literature. Our approach provides an exploration bonus that rewards the agent for gathering information about the state that is relevant for completing the task. While this requires the agent to know what this information is during training, it can obtained in several ways: in the most general case, offpolicy algorithms can leverage knowledge about the entire trajectory to determine such information in hindsight, but the user can also provide prior knowledge (e.g., privileged information) to help inform the training process. Through experiments in several partially-observed environments, we find that our approach is competitive with prior methods when minimal exploration is needed, but substantially outperforms them when more complex strategies are required. Our algorithm also shows the ability to learn without any privileged information, by reasoning about the entire trajectory in hindsight and and effectively using any information it reveals about the hidden state.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext cd81130f-0f04-4c16-9e61-d7632900abc7Cited by top-tier papers2
- RefactorBench: Evaluating Stateful Reasoning in Language Agents Through CodeDhruv Gautam, Spandan Garg, Jinu Jang, Neel Sundaresan et al.ICLR 2025
- Adaptive Exploration for Latent-State BanditsJikai Jin, Kenneth Hung, Sanath Kumar Krishnamurthy, Baoyi Shi et al.KDD 2026
Builds on16
- Stochastic Latent Actor-Critic: Deep Reinforcement Learning with a Latent Variable ModelAlex X. Lee, Anusha Nagabandi, Pieter Abbeel, Sergey LevineNeurIPS 2020 · 437 citations
- Decoupling Exploration and Exploitation for Meta-Reinforcement Learning without SacrificesEvan Zheran Liu, Aditi Raghunathan, Percy Liang, Chelsea FinnICML 2021 · 80 citations
- Variational Recurrent Models for Solving Partially Observable Control TasksDongqi Han, Kenji Doya, Jun TaniICLR 2020 · 75 citations
- Counterfactual Credit Assignment in Model-Free Reinforcement LearningThomas Mesnard, Theophane Weber, Fabio Viola, Shantanu Thakoor et al.ICML 2021 · 70 citations
- Bridging the Imitation Gap by Adaptive InsubordinationLuca Weihs, Unnat Jain, Iou-Jen Liu, Jordi Salvador et al.NeurIPS 2021 · 53 citations
Related papers
- Information Particle Filter Tree: An Online Algorithm for POMDPs with Belief-Based Rewards on Continuous DomainsJohannes Fischer, Ömer Sahin TasICML 2020 · 42 citations
- Learning Belief Representations for Partially Observable Deep RLAndrew Wang, Andrew C. Li, Toryn Q. Klassen, Rodrigo Toro Icarte et al.ICML 2023 · 21 citations
- Learning in POMDPs is Sample-Efficient with Hindsight ObservabilityJonathan Lee, Alekh Agarwal, Christoph Dann, Tong ZhangICML 2023 · 25 citations
- Belief-Dependent Macro-Action Discovery in POMDPs using the Value of InformationGenevieve Flaspohler, Nicholas Roy, John W. Fisher IIINeurIPS 2020 · 13 citations
- Revelations: A Decidable Class of POMDPs with Omega-Regular ObjectivesMarius Belly, Nathanaël Fijalkow, Hugo Gimbert, Florian Horn et al.AAAI 2025 · 5 citations
