A Consciousness-Inspired Planning Agent for Model-Based Reinforcement Learning
Mingde Zhao, Zhen Liu, Sitao Luan, Shuyuan Zhang, Doina Precup, Yoshua Bengio
Abstract
We present an end-to-end, model-based deep reinforcement learning agent which dynamically attends to relevant parts of its state during planning. The agent uses a bottleneck mechanism over a set-based representation to force the number of entities to which the agent attends at each planning step to be small. In experiments, we investigate the bottleneck mechanism with several sets of customized environments featuring different challenges. We consistently observe that the design allows the planning agents to generalize their learned task-solving abilities in compatible unseen environments by attending to the relevant objects, leading to better out-of-distribution generalization performance.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext ec66ca91-2aca-4afc-a965-7b59b9e3c942Cited by top-tier papers9
- Dynamic Model Predictive Shielding for Provably Safe Reinforcement LearningArko Banerjee, Kia Rahmani, Joydeep Biswas, Isil DilligNeurIPS 2024 · 24 citations
- SPARTAN: A Sparse Transformer World Model Attending to What MattersAnson Lei, Bernhard Schölkopf, Ingmar PosnerNeurIPS 2025 · 12 citations
- Translate Policy to Language: Flow Matching Generated Rewards for LLM ExplanationsXinyi Yang, Liang Zeng, Heng Dong, Chao Yu et al.ICLR 2026 · 6 citations
- Reusable Slotwise MechanismsBailey Trang Nguyen, Amin Mansouri, Kanika Madan, Khuong Nguyen et al.NeurIPS 2023 · 6 citations
- Consciousness-Inspired Spatio-Temporal Abstractions for Better Generalization in Reinforcement LearningHarry Zhao, Safa Alver, Harm van Seijen, Romain Laroche et al.ICLR 2024 · 5 citations
Builds on5
- Mastering Atari with Discrete World ModelsDanijar Hafner, Timothy P. Lillicrap, Mohammad Norouzi, Jimmy BaICLR 2021 · 1,170 citations
- Model Based Reinforcement Learning for AtariLukasz Kaiser, Mohammad Babaeizadeh, Piotr Milos, Blazej Osinski et al.ICLR 2020 · 969 citations
- On the role of planning in model-based deep reinforcement learningJessica B. Hamrick, Abram L. Friesen, Feryal M. P. Behbahani, Arthur Guez et al.ICLR 2021 · 77 citations
- Temporally Abstract Partial ModelsKhimya Khetarpal, Zafarali Ahmed, Gheorghe Comanici, Doina PrecupNeurIPS 2021 · 17 citations
- Refactoring Policy for Compositional Generalizability using Self-Supervised Object ProposalsTongzhou Mu, Jiayuan Gu, Zhiwei Jia, Hao Tang et al.NeurIPS 2020 · 13 citations
Related papers
- Fast And Slow Learning Of Recurrent Independent MechanismsKanika Madan, Nan Rosemary Ke, Anirudh Goyal, Bernhard Schölkopf et al.ICLR 2021 · 41 citations
- Discrete Compositional Representations as an Abstraction for Goal Conditioned Reinforcement LearningRiashat Islam, Hongyu Zang, Anirudh Goyal, Alex Lamb et al.NeurIPS 2022 · 10 citations
- DRIBO: Robust Deep Reinforcement Learning via Multi-View Information BottleneckJiameng Fan, Wenchao LiICML 2022 · 49 citations
- SOLD: Slot Object-Centric Latent Dynamics Models for Relational Manipulation Learning from PixelsMalte Mosbach, Jan Niklas Ewertz, Angel Villar-Corrales, Sven BehnkeICML 2025
- 3D-aware Disentangled Representation for Compositional Reinforcement LearningSungbin Mun, Younghwan Lee, Cheolhui MIn, Mineui Hong et al.ICLR 2026
