Discriminative Particle Filter Reinforcement Learning for Complex Partial observations
Xiao Ma, Péter Karkus, David Hsu, Wee Sun Lee, Nan Ye
Abstract
Deep reinforcement learning has succeeded in sophisticated games such as Atari, Go, etc. Real-world decision making, however, often requires reasoning with partial information extracted from complex visual observations. This paper presents Discriminative Particle Filter Reinforcement Learning (DPFRL), a new reinforcement learning framework for partial and complex observations. DPFRL encodes a differentiable particle filter with learned transition and observation models in a neural network, which allows for reasoning with partial observations over multiple time steps. While a standard particle filter relies on a generative observation model, DPFRL learns a discriminatively parameterized model that is training directly for decision making. We show that the discriminative parameterization results in significantly improved performance, especially for tasks with complex visual observations, because it circumvents the difficulty of modelling observations explicitly. In most cases, DPFRL outperforms state-of-the-art POMDP RL models in Flickering Atari Games, an existing POMDP RL benchmark, and in Natural Flickering Atari Games, a new, more challenging POMDP RL benchmark that we introduce. We further show that DPFRL performs well for visual navigation with real-world data.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext afe22b15-a741-40ce-8ef7-ca080235e979Cited by top-tier papers6
- Differentiable Particle Filtering via Entropy-Regularized Optimal TransportAdrien Corenflos, James Thornton, George Deligiannidis, Arnaud DoucetICML 2021 · 91 citations
- Online Variational Filtering and Parameter LearningAndrew Campbell, Yuyang Shi, Thomas Rainforth, Arnaud DoucetNeurIPS 2021 · 30 citations
- PF-GNN: Differentiable particle filtering based approximation of universal graph representationsMohammed Haroon Dupty, Yanfei Dong, Wee Sun LeeICLR 2022 · 14 citations
- Sequential Monte Carlo for Policy Optimization in Continuous POMDPsHany Abdulsamad, Sahel Mohammad Iqbal, Simo SärkkäNeurIPS 2025 · 4 citations
- Set-membership Belief State-based Reinforcement Learning for POMDPsWei Wei, Lijun Zhang, Lin Li, Huizhong Song et al.ICML 2023 · 2 citations
Builds on2
Related papers
- Learning Belief Representations for Partially Observable Deep RLAndrew Wang, Andrew C. Li, Toryn Q. Klassen, Rodrigo Toro Icarte et al.ICML 2023 · 21 citations
- Diffusion for World Modeling: Visual Details Matter in AtariEloi Alonso, Adam Jelley, Vincent Micheli, Anssi Kanervisto et al.NeurIPS 2024 · 359 citations
- Provable Representation with Efficient Planning for Partially Observable Reinforcement LearningHongming Zhang, Tongzheng Ren, Chenjun Xiao, Dale Schuurmans et al.ICML 2024 · 9 citations
- Flow-based Recurrent Belief State Learning for POMDPsXiaoyu Chen, Yao Mark Mu, Ping Luo, Shengbo Li et al.ICML 2022 · 26 citations
- Differentiable and Stable Long-Range Tracking of Multiple Posterior ModesAli Younis, Erik B. SudderthNeurIPS 2023 · 7 citations
