Discriminative Particle Filter Reinforcement Learning for Complex Partial observations
Xiao Ma, Péter Karkus, David Hsu, Wee Sun Lee, Nan Ye
摘要
Deep reinforcement learning has succeeded in sophisticated games such as Atari, Go, etc. Real-world decision making, however, often requires reasoning with partial information extracted from complex visual observations. This paper presents Discriminative Particle Filter Reinforcement Learning (DPFRL), a new reinforcement learning framework for partial and complex observations. DPFRL encodes a differentiable particle filter with learned transition and observation models in a neural network, which allows for reasoning with partial observations over multiple time steps. While a standard particle filter relies on a generative observation model, DPFRL learns a discriminatively parameterized model that is training directly for decision making. We show that the discriminative parameterization results in significantly improved performance, especially for tasks with complex visual observations, because it circumvents the difficulty of modelling observations explicitly. In most cases, DPFRL outperforms state-of-the-art POMDP RL models in Flickering Atari Games, an existing POMDP RL benchmark, and in Natural Flickering Atari Games, a new, more challenging POMDP RL benchmark that we introduce. We further show that DPFRL performs well for visual navigation with real-world data.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- Differentiable Particle Filtering via Entropy-Regularized Optimal TransportAdrien Corenflos, James Thornton, George Deligiannidis, Arnaud DoucetICML 2021 · 被引用 91 次
- Online Variational Filtering and Parameter LearningAndrew Campbell, Yuyang Shi, Thomas Rainforth, Arnaud DoucetNeurIPS 2021 · 被引用 30 次
- PF-GNN: Differentiable particle filtering based approximation of universal graph representationsMohammed Haroon Dupty, Yanfei Dong, Wee Sun LeeICLR 2022 · 被引用 14 次
- Sequential Monte Carlo for Policy Optimization in Continuous POMDPsHany Abdulsamad, Sahel Mohammad Iqbal, Simo SärkkäNeurIPS 2025 · 被引用 4 次
- Set-membership Belief State-based Reinforcement Learning for POMDPsWei Wei, Lijun Zhang, Lin Li, Huizhong Song 等ICML 2023 · 被引用 2 次
它引用的顶会 Paper2
相关 Paper
- Learning Belief Representations for Partially Observable Deep RLAndrew Wang, Andrew C. Li, Toryn Q. Klassen, Rodrigo Toro Icarte 等ICML 2023 · 被引用 21 次
- Diffusion for World Modeling: Visual Details Matter in AtariEloi Alonso, Adam Jelley, Vincent Micheli, Anssi Kanervisto 等NeurIPS 2024 · 被引用 359 次
- Provable Representation with Efficient Planning for Partially Observable Reinforcement LearningHongming Zhang, Tongzheng Ren, Chenjun Xiao, Dale Schuurmans 等ICML 2024 · 被引用 9 次
- Flow-based Recurrent Belief State Learning for POMDPsXiaoyu Chen, Yao Mark Mu, Ping Luo, Shengbo Li 等ICML 2022 · 被引用 26 次
- Differentiable and Stable Long-Range Tracking of Multiple Posterior ModesAli Younis, Erik B. SudderthNeurIPS 2023 · 被引用 7 次
