Learning to Explore in POMDPs with Informational Rewards
Annie Xie, Logan M. Bhamidipaty, Evan Zheran Liu, Joey Hong, Sergey Levine, Chelsea Finn
摘要
Standard exploration methods typically rely on random coverage of the state space or coveragepromoting exploration bonuses. However, in partially observed settings, the biggest exploration challenge is often posed by the need to discover information-gathering strategies-e.g., an agent that has to navigate to a location in traffic might learn to first check traffic conditions and then choose a route. In this work, we design a POMDP agent that gathers information about the hidden state, using ideas from the meta-exploration literature. Our approach provides an exploration bonus that rewards the agent for gathering information about the state that is relevant for completing the task. While this requires the agent to know what this information is during training, it can obtained in several ways: in the most general case, offpolicy algorithms can leverage knowledge about the entire trajectory to determine such information in hindsight, but the user can also provide prior knowledge (e.g., privileged information) to help inform the training process. Through experiments in several partially-observed environments, we find that our approach is competitive with prior methods when minimal exploration is needed, but substantially outperforms them when more complex strategies are required. Our algorithm also shows the ability to learn without any privileged information, by reasoning about the entire trajectory in hindsight and and effectively using any information it reveals about the hidden state.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- RefactorBench: Evaluating Stateful Reasoning in Language Agents Through CodeDhruv Gautam, Spandan Garg, Jinu Jang, Neel Sundaresan 等ICLR 2025
- Adaptive Exploration for Latent-State BanditsJikai Jin, Kenneth Hung, Sanath Kumar Krishnamurthy, Baoyi Shi 等KDD 2026
它引用的顶会 Paper16
- Stochastic Latent Actor-Critic: Deep Reinforcement Learning with a Latent Variable ModelAlex X. Lee, Anusha Nagabandi, Pieter Abbeel, Sergey LevineNeurIPS 2020 · 被引用 437 次
- Decoupling Exploration and Exploitation for Meta-Reinforcement Learning without SacrificesEvan Zheran Liu, Aditi Raghunathan, Percy Liang, Chelsea FinnICML 2021 · 被引用 80 次
- Variational Recurrent Models for Solving Partially Observable Control TasksDongqi Han, Kenji Doya, Jun TaniICLR 2020 · 被引用 75 次
- Counterfactual Credit Assignment in Model-Free Reinforcement LearningThomas Mesnard, Theophane Weber, Fabio Viola, Shantanu Thakoor 等ICML 2021 · 被引用 70 次
- Bridging the Imitation Gap by Adaptive InsubordinationLuca Weihs, Unnat Jain, Iou-Jen Liu, Jordi Salvador 等NeurIPS 2021 · 被引用 53 次
相关 Paper
- Information Particle Filter Tree: An Online Algorithm for POMDPs with Belief-Based Rewards on Continuous DomainsJohannes Fischer, Ömer Sahin TasICML 2020 · 被引用 42 次
- Learning Belief Representations for Partially Observable Deep RLAndrew Wang, Andrew C. Li, Toryn Q. Klassen, Rodrigo Toro Icarte 等ICML 2023 · 被引用 21 次
- Learning in POMDPs is Sample-Efficient with Hindsight ObservabilityJonathan Lee, Alekh Agarwal, Christoph Dann, Tong ZhangICML 2023 · 被引用 25 次
- Belief-Dependent Macro-Action Discovery in POMDPs using the Value of InformationGenevieve Flaspohler, Nicholas Roy, John W. Fisher IIINeurIPS 2020 · 被引用 13 次
- Revelations: A Decidable Class of POMDPs with Omega-Regular ObjectivesMarius Belly, Nathanaël Fijalkow, Hugo Gimbert, Florian Horn 等AAAI 2025 · 被引用 5 次
