Learning Interactive World Model for Object-Centric Reinforcement Learning
Fan Feng, Phillip Lippe, Sara Magliacane
Abstract
Agents that understand objects and their interactions can learn policies that are more robust and transferable. However, most object-centric RL methods factor state by individual objects while leaving interactions implicit. We introduce the Factored Interactive Object-Centric World Model (FIOC-WM), a unified framework that learns structured representations of both objects and their interactions within a world model. FIOC-WM captures environment dynamics with disentangled and modular representations of object interactions, improving sample efficiency and generalization for policy learning. Concretely, FIOC-WM first learns object-centric latents and an interaction structure directly from pixels, leveraging pre-trained vision encoders. The learned world model then decomposes tasks into composable interaction primitives, and a hierarchical policy is trained on top: a high level selects the type and order of interactions, while a low level executes them. On simulated robotic and embodied-AI benchmarks, FIOC-WM improves policy-learning sample efficiency and generalization over world-model baselines, indicating that explicit, modular interaction learning is crucial for robust control 2 .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e08247bb-403c-430a-b31b-a2f4290a0233Cited by top-tier papers3
- Causal-JEPA: Learning World Models through Object-Level Latent MaskingHeejeong Nam, Quentin Le Lidec, Lucas Maes, Yann LeCun et al.ICML 2026 · 7 citations
- Ada-Diffuser: Latent-Aware Adaptive Diffusion for Decision-MakingFan Feng, Selena Ge, Minghao Fu, Zijian Li et al.ICLR 2026 · 3 citations
- Relational Structural Causal ModelsAdiba Ejaz, Elias BareinboimICML 2026
Builds on50
- Dream to Control: Learning Behaviors by Latent ImaginationDanijar Hafner, Timothy P. Lillicrap, Jimmy Ba, Mohammad NorouziICLR 2020 · 1,852 citations
- Object-Centric Learning with Slot AttentionFrancesco Locatello, Dirk Weissenborn, Thomas Unterthiner, Aravindh Mahendran et al.NeurIPS 2020 · 1,275 citations
- Mastering Atari with Discrete World ModelsDanijar Hafner, Timothy P. Lillicrap, Mohammad Norouzi, Jimmy BaICLR 2021 · 1,170 citations
- Learning Universal Policies via Text-Guided Video GenerationYilun Du, Sherry Yang, Bo Dai, Hanjun Dai et al.NeurIPS 2023 · 742 citations
- Learning Interactive Real-World SimulatorsSherry Yang, Yilun Du, Seyed Kamyar Seyed Ghasemipour, Jonathan Tompson et al.ICLR 2024 · 399 citations
Related papers
- Learning Dynamic Attribute-factored World Models for Efficient Multi-object Reinforcement LearningFan Feng, Sara MagliacaneNeurIPS 2023 · 17 citations
- Dyn-O: Building Structured World Models with Object-Centric RepresentationsZizhao Wang, Kaixin Wang, Li Zhao, Peter Stone et al.NeurIPS 2025 · 15 citations
- Object-Centric World Models for Causality-Aware Reinforcement LearningYosuke Nishimoto, Takashi MatsubaraAAAI 2026 · 2 citations
- Learning Task-Sufficient World Models by Synergizing Agentic Exploration and Structured ModelingFan Feng, Yujia Zheng, Minghao Fu, Yongqiang Chen et al.ICML 2026
- Learning Disentangled Multi-Agent World Model for Decentralized ControlDi Xue, Jing Jiang, Shaowei Zhang, Wenhao Guo et al.ICML 2026 · 2 citations
