The Wasserstein Believer: Learning Belief Updates for Partially Observable Environments through Reliable Latent Space Models
Raphaël Avalos, Florent Delgrange, Ann Nowé, Guillermo A. Pérez, Diederik M. Roijers
摘要
Partially Observable Markov Decision Processes (POMDPs) are used to model environments where the full state cannot be perceived by an agent. As such the agent needs to reason taking into account the past observations and actions. However, simply remembering the full history is generally intractable due to the exponential growth in the history space. Maintaining a probability distribution that models the belief over what the true state is can be used as a sufficient statistic of the history, but its computation requires access to the model of the environment and is often intractable. While SOTA algorithms use Recurrent Neural Networks to compress the observation-action history aiming to learn a sufficient statistic, they lack guarantees of success and can lead to sub-optimal policies. To overcome this, we propose the Wasserstein Belief Updater, an RL algorithm that learns a latent model of the POMDP and an approximation of the belief update. Our approach comes with theoretical guarantees on the quality of our approximation ensuring that our outputted beliefs allow for learning the optimal value function.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Provable Partially Observable Reinforcement Learning with Privileged InformationYang Cai, Xiangyu Liu, Argyris Oikonomou, Kaiqing ZhangNeurIPS 2024 · 被引用 22 次
- Deep SPI: Safe Policy Improvement via World ModelsFlorent Delgrange, Raphaël Avalos, Willem RöpkeICLR 2026 · 被引用 4 次
- Informed Asymmetric Actor-Critic: Leveraging Privileged Signals Beyond Full-State AccessDaniel Ebi, Damien Ernst, Klemens Böhm, Gaspard LambrechtsICML 2026 · 被引用 1 次
- PIGDreamer: Privileged Information Guided World Models for Safe Partially Observable Reinforcement LearningDongchi Huang, Jiaqi Wang, Yang Li, Chunhe Xia 等ICML 2025
它引用的顶会 Paper11
- Dream to Control: Learning Behaviors by Latent ImaginationDanijar Hafner, Timothy P. Lillicrap, Jimmy Ba, Mohammad NorouziICLR 2020 · 被引用 1,852 次
- Mastering Atari with Discrete World ModelsDanijar Hafner, Timothy P. Lillicrap, Mohammad Norouzi, Jimmy BaICLR 2021 · 被引用 1,170 次
- Stochastic Latent Actor-Critic: Deep Reinforcement Learning with a Latent Variable ModelAlex X. Lee, Anusha Nagabandi, Pieter Abbeel, Sergey LevineNeurIPS 2020 · 被引用 437 次
- Particle Filter Recurrent Neural NetworksXiao Ma, Péter Karkus, David Hsu, Wee Sun LeeAAAI 2020 · 被引用 94 次
- Steady State Analysis of Episodic Reinforcement LearningBojun HuangNeurIPS 2020 · 被引用 29 次
相关 Paper
- Flow-based Recurrent Belief State Learning for POMDPsXiaoyu Chen, Yao Mark Mu, Ping Luo, Shengbo Li 等ICML 2022 · 被引用 26 次
- Learning Belief Representations for Partially Observable Deep RLAndrew Wang, Andrew C. Li, Toryn Q. Klassen, Rodrigo Toro Icarte 等ICML 2023 · 被引用 21 次
- SVQN: Sequential Variational Soft Q-Learning NetworksShiyu Huang, Hang Su, Jun Zhu, Ting ChenICLR 2020 · 被引用 19 次
- Probing in the Dark: State Entropy Maximization for POMDPsYonatan Ashlag, Mirco Mutti, Aviv Tamar, Kfir LevyICLR 2026
- Scalable Policy-Based RL Algorithms for POMDPsAmeya Anjarlekar, S. Rasoul Etesami, R. SrikantNeurIPS 2025 · 被引用 6 次
