A Consciousness-Inspired Planning Agent for Model-Based Reinforcement Learning
Mingde Zhao, Zhen Liu, Sitao Luan, Shuyuan Zhang, Doina Precup, Yoshua Bengio
摘要
We present an end-to-end, model-based deep reinforcement learning agent which dynamically attends to relevant parts of its state during planning. The agent uses a bottleneck mechanism over a set-based representation to force the number of entities to which the agent attends at each planning step to be small. In experiments, we investigate the bottleneck mechanism with several sets of customized environments featuring different challenges. We consistently observe that the design allows the planning agents to generalize their learned task-solving abilities in compatible unseen environments by attending to the relevant objects, leading to better out-of-distribution generalization performance.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper9
- Dynamic Model Predictive Shielding for Provably Safe Reinforcement LearningArko Banerjee, Kia Rahmani, Joydeep Biswas, Isil DilligNeurIPS 2024 · 被引用 24 次
- SPARTAN: A Sparse Transformer World Model Attending to What MattersAnson Lei, Bernhard Schölkopf, Ingmar PosnerNeurIPS 2025 · 被引用 12 次
- Translate Policy to Language: Flow Matching Generated Rewards for LLM ExplanationsXinyi Yang, Liang Zeng, Heng Dong, Chao Yu 等ICLR 2026 · 被引用 6 次
- Reusable Slotwise MechanismsBailey Trang Nguyen, Amin Mansouri, Kanika Madan, Khuong Nguyen 等NeurIPS 2023 · 被引用 6 次
- Consciousness-Inspired Spatio-Temporal Abstractions for Better Generalization in Reinforcement LearningHarry Zhao, Safa Alver, Harm van Seijen, Romain Laroche 等ICLR 2024 · 被引用 5 次
它引用的顶会 Paper5
- Mastering Atari with Discrete World ModelsDanijar Hafner, Timothy P. Lillicrap, Mohammad Norouzi, Jimmy BaICLR 2021 · 被引用 1,170 次
- Model Based Reinforcement Learning for AtariLukasz Kaiser, Mohammad Babaeizadeh, Piotr Milos, Blazej Osinski 等ICLR 2020 · 被引用 969 次
- On the role of planning in model-based deep reinforcement learningJessica B. Hamrick, Abram L. Friesen, Feryal M. P. Behbahani, Arthur Guez 等ICLR 2021 · 被引用 77 次
- Temporally Abstract Partial ModelsKhimya Khetarpal, Zafarali Ahmed, Gheorghe Comanici, Doina PrecupNeurIPS 2021 · 被引用 17 次
- Refactoring Policy for Compositional Generalizability using Self-Supervised Object ProposalsTongzhou Mu, Jiayuan Gu, Zhiwei Jia, Hao Tang 等NeurIPS 2020 · 被引用 13 次
相关 Paper
- Fast And Slow Learning Of Recurrent Independent MechanismsKanika Madan, Nan Rosemary Ke, Anirudh Goyal, Bernhard Schölkopf 等ICLR 2021 · 被引用 41 次
- Discrete Compositional Representations as an Abstraction for Goal Conditioned Reinforcement LearningRiashat Islam, Hongyu Zang, Anirudh Goyal, Alex Lamb 等NeurIPS 2022 · 被引用 10 次
- DRIBO: Robust Deep Reinforcement Learning via Multi-View Information BottleneckJiameng Fan, Wenchao LiICML 2022 · 被引用 49 次
- SOLD: Slot Object-Centric Latent Dynamics Models for Relational Manipulation Learning from PixelsMalte Mosbach, Jan Niklas Ewertz, Angel Villar-Corrales, Sven BehnkeICML 2025
- 3D-aware Disentangled Representation for Compositional Reinforcement LearningSungbin Mun, Younghwan Lee, Cheolhui MIn, Mineui Hong 等ICLR 2026
