Interaction-Based Disentanglement of Entities for Object-Centric World Models
Akihiro Nakano, Masahiro Suzuki, Yutaka Matsuo
摘要
Perceiving the world compositionally in terms of space and time is essential to understanding object dynamics and solving downstream tasks. Object-centric learning using generative models has improved in its ability to learn distinct representations of individual objects and predict their interactions, and how to utilize the learned representations to solve untrained, downstream tasks is a focal question. However, as models struggle to predict object interactions and track the objects accurately, especially for unseen configurations, using object-centric representations in downstream tasks is still a challenge. This paper proposes STEDIE, a new model that disentangles object representations, based on interactions, into interaction-relevant relational features and interaction-irrelevant global features without supervision. Empirical evaluation shows that the proposed model factorizes global features, unaffected by interactions from relational features that are necessary to predict outcome of interactions. We also show that STEDIE achieves better performance in planning tasks and understanding causal relationships. In both tasks, our model not only achieves better performance in terms of reconstruction ability but also utilizes the disentangled representations to solve the tasks in a structured manner.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
引用它的顶会 Paper7
- Learning Dynamic Attribute-factored World Models for Efficient Multi-object Reinforcement LearningFan Feng, Sara MagliacaneNeurIPS 2023 · 被引用 17 次
- Dyn-O: Building Structured World Models with Object-Centric RepresentationsZizhao Wang, Kaixin Wang, Li Zhao, Peter Stone 等NeurIPS 2025 · 被引用 15 次
- Learning Interactive World Model for Object-Centric Reinforcement LearningFan Feng, Phillip Lippe, Sara MagliacaneNeurIPS 2025 · 被引用 13 次
- Latent Particle World Models: Self-supervised Object-centric Stochastic Dynamics ModelingTal Daniel, Carl Qi, Dan Haramati, Amir Zadeh 等ICLR 2026 · 被引用 12 次
- United We Stand: Towards End-to-End Log-based Fault Diagnosis via Interactive Multi-Task LearningMinghua He, Chiming Duan, Pei Xiao, Tong Jia 等ASE 2025 · 被引用 1 次
相关 Paper
- GATSBI: Generative Agent-Centric Spatio-Temporal Object InteractionCheol-Hui Min, Jinseok Bae, Junho Lee, Young Min KimCVPR 2021
- Robust and Controllable Object-Centric Learning through Energy-based ModelsRuixiang Zhang, Tong Che, Boris Ivanovic, Renhao Wang 等ICLR 2023
- SlotFormer: Unsupervised Visual Dynamics Simulation with Object-Centric ModelsZiyi Wu, Nikita Dvornik, Klaus Greff, Thomas Kipf 等ICLR 2023 · 被引用 10 次
- VDSM: Unsupervised Video Disentanglement With State-Space Modeling and Deep Mixtures of ExpertsMatthew J. Vowels, Necati Cihan Camgöz, Richard BowdenCVPR 2021
- RELATE: Physically Plausible Multi-Object Scene Synthesis Using Structured Latent SpacesSébastien Ehrhardt, Oliver Groth, Áron Monszpart, Martin Engelcke 等NeurIPS 2020 · 被引用 61 次
