Information-based Value Iteration Networks for Decision Making Under Uncertainty
Cynthia Chen, Samantha Johnson, Cindy Poo, Michael A Buice, Koosha Khalvati
摘要
Deep neural networks that incorporate classic reinforcement learning methods, such as value iteration, into their structure significantly outperform randomly structured networks in learning and generalization. These networks, however, are mostly limited to environments with no or very low uncertainty and do not extend well to partially observable environments. In this paper, we propose a new planning module architecture, the VIN (Value Iteration with Value of Information Network), that learns to act in novel environments with high perceptual ambiguity. This architecture over-emphasizes reducing uncertainty before exploiting the reward. VIN can also utilize factorization in environments with mixed observability to decrease the computational complexity of calculating the policy and to facilitate learning. Tested on a range of grid-based navigation tasks, each containing various types of environments with different degrees of observability, our network outperforms other deep architectures. Moreover, VIN generates interpretable cognitive maps highlighting both rewarding and informative locations. These maps highlight the key states the agent must visit to achieve its goal.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper4
- Decision Transformer: Reinforcement Learning via Sequence ModelingLili Chen, Kevin Lu, Aravind Rajeswaran, Kimin Lee 等NeurIPS 2021 · 被引用 2,557 次
- Recurrent Model-Free RL Can Be a Strong Baseline for Many POMDPsTianwei Ni, Benjamin Eysenbach, Ruslan SalakhutdinovICML 2022 · 被引用 162 次
- Towards real-world navigation with deep differentiable plannersShu Ishida, João F. HenriquesCVPR 2022 · 被引用 6 次
- Universal Value Iteration Networks: When Spatially-Invariant Is Not UniversalLi Zhang, Xin Li, Sen Chen, Hongyu Zang 等AAAI 2020 · 被引用 5 次
相关 Paper
- Highway Value Iteration NetworksYuhui Wang, Weida Li, Francesco Faccio, Qingyuan Wu 等ICML 2024 · 被引用 3 次
- Variational Recurrent Models for Solving Partially Observable Control TasksDongqi Han, Kenji Doya, Jun TaniICLR 2020 · 被引用 75 次
- SVQN: Sequential Variational Soft Q-Learning NetworksShiyu Huang, Hang Su, Jun Zhu, Ting ChenICLR 2020 · 被引用 19 次
- Versatile Navigation Under Partial Observability via Value-Guided Diffusion PolicyGengyu Zhang, Hao Tang, Yan YanCVPR 2024 · 被引用 2 次
- Predictive Coding Enhances Meta-RL To Achieve Interpretable Bayes-Optimal Belief Representation Under Partial ObservabilityPo-Chen Kuo, Han Hou, Will Dabney, Edgar Y. WalkerNeurIPS 2025
