Deep Recurrent Belief Propagation Network for POMDPs
Yuhui Wang, Xiaoyang Tan
Abstract
In many real-world sequential decision-making tasks, especially in continuous control like robotic control, it is rare that the observations are perfect, that is, the sensory data could be incomplete, noisy or even dynamically polluted due to the unexpected malfunctions or intrinsic low quality of the sensors. Previous methods handle these issues in the framework of POMDPs and are either deterministic by feature memorization or stochastic by belief inference. In this paper, we present a new method that lies somewhere in the middle of the spectrum of research methodology identified above and combines the strength of both approaches. In particular, the proposed method, named Deep Recurrent Belief Propagation Network (DRBPN), takes a hybrid style belief updating procedure − an RNN-type feature extraction step followed by an analytical belief inference, significantly reducing the computational cost while faithfully capturing the complex dynamics and maintaining the necessary uncertainty for generalization. The effectiveness of the proposed method is verified on a collection of benchmark tasks, showing that our approach outperforms several state-of-the-art methods under various challenging scenarios.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e2bc8a3c-42b4-4a5e-bbe4-1ce51390ab99Cited by top-tier papers6
- Learning Belief Representations for Partially Observable Deep RLAndrew Wang, Andrew C. Li, Toryn Q. Klassen, Rodrigo Toro Icarte et al.ICML 2023 · 21 citations
- The Wasserstein Believer: Learning Belief Updates for Partially Observable Environments through Reliable Latent Space ModelsRaphaël Avalos, Florent Delgrange, Ann Nowé, Guillermo A. Pérez et al.ICLR 2024 · 10 citations
- Structure Learning-Based Task Decomposition for Reinforcement Learning in Non-stationary EnvironmentsHonguk Woo, Gwangpyo Yoo, Minjong YooAAAI 2022 · 5 citations
- Dual Critic Reinforcement Learning under Partial ObservabilityJinqiu Li, Enmin Zhao, Tong Wei, Junliang Xing et al.NeurIPS 2024 · 3 citations
- Set-membership Belief State-based Reinforcement Learning for POMDPsWei Wei, Lijun Zhang, Lin Li, Huizhong Song et al.ICML 2023 · 2 citations
Related papers
- Flow-based Recurrent Belief State Learning for POMDPsXiaoyu Chen, Yao Mark Mu, Ping Luo, Shengbo Li et al.ICML 2022 · 26 citations
- SVQN: Sequential Variational Soft Q-Learning NetworksShiyu Huang, Hang Su, Jun Zhu, Ting ChenICLR 2020 · 19 citations
- ODE-based Recurrent Model-free Reinforcement Learning for POMDPsXuanle Zhao, Duzhen Zhang, Liyuan Han, Tielin Zhang et al.NeurIPS 2023 · 18 citations
- Variational Recurrent Models for Solving Partially Observable Control TasksDongqi Han, Kenji Doya, Jun TaniICLR 2020 · 75 citations
- Sample-efficient and Scalable Exploration in Continuous-Time RLKlemens Iten, Lenart Treven, Bhavya Sukhija, Florian Dörfler et al.ICLR 2026 · 3 citations
