Entity-Centric Reinforcement Learning for Object Manipulation from Pixels
Dan Haramati, Tal Daniel, Aviv Tamar
Abstract
Manipulating objects is a hallmark of human intelligence, and an important task in domains such as robotics. In principle, Reinforcement Learning (RL) offers a general approach to learn object manipulation. In practice, however, domains with more than a few objects are difficult for RL agents due to the curse of dimensionality, especially when learning from raw image observations. In this work we propose a structured approach for visual RL that is suitable for representing multiple objects and their interaction, and use it to learn goal-conditioned manipulation of several objects. Key to our method is the ability to handle goals with dependencies between the objects (e.g., moving objects in a certain order). We further relate our architecture to the generalization capability of the trained agent, based on a theoretical result for compositional generalization, and demonstrate agents that learn with 3 objects but generalize to similar tasks with over 10 objects. Videos and code are available on the project website: https://sites.google.com/view/entity-centric-rl
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers20
- Learning Interactive World Model for Object-Centric Reinforcement LearningFan Feng, Phillip Lippe, Sara MagliacaneNeurIPS 2025 · 13 citations
- Latent Particle World Models: Self-supervised Object-centric Stochastic Dynamics ModelingTal Daniel, Carl Qi, Dan Haramati, Amir Zadeh et al.ICLR 2026 · 12 citations
- MetaSlot: Break Through the Fixed Number of Slots in Object-Centric LearningHongjia Liu, Rongzhen Zhao, Haohan Chen, Joni PajarinenNeurIPS 2025 · 12 citations
- Object-Centric World Models from Few-Shot Annotations for Sample-Efficient Reinforcement LearningWeipu Zhang, Adam Jelley, Trevor McInroe, Amos J. Storkey et al.ICLR 2026 · 9 citations
- Hierarchical Entity-centric Reinforcement Learning with Factored Subgoal DiffusionDan Haramati, Carl Qi, Tal Daniel, Amy Zhang et al.ICLR 2026 · 7 citations
Builds on18
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Object-Centric Learning with Slot AttentionFrancesco Locatello, Dirk Weissenborn, Thomas Unterthiner, Aravindh Mahendran et al.NeurIPS 2020 · 1,275 citations
- Improving Sample Efficiency in Model-Free Reinforcement Learning from ImagesDenis Yarats, Amy Zhang, Ilya Kostrikov, Brandon Amos et al.AAAI 2021 · 506 citations
- Skew-Fit: State-Covering Self-Supervised Reinforcement LearningVitchyr Pong, Murtaza Dalal, Steven Lin, Ashvin Nair et al.ICML 2020 · 303 citations
- Illiterate DALL-E Learns to ComposeGautam Singh, Fei Deng, Sungjin AhnICLR 2022 · 182 citations
Related papers
- Hierarchical Abstraction for Combinatorial Generalization in Object RearrangementMichael Chang, Alyssa L. Dayan, Franziska Meier, Thomas L. Griffiths et al.ICLR 2023
- 3D-aware Disentangled Representation for Compositional Reinforcement LearningSungbin Mun, Younghwan Lee, Cheolhui MIn, Mineui Hong et al.ICLR 2026
- Learning Object-Centric Motion Priors from Human for Robotic Dexterous ManipulationZhengdong Hong, Guofeng ZhangAAAI 2026
- Multi-fingered Hand Grasps with Visuo-Tactile Fusion via Multi-Agent Deep Reinforcement LearningPeida Jia, Xuanheng Li, Tianqiang Zhu, Rina Wu et al.AAAI 2025 · 3 citations
- Learning Dynamic Attribute-factored World Models for Efficient Multi-object Reinforcement LearningFan Feng, Sara MagliacaneNeurIPS 2023 · 17 citations
