3D-aware Disentangled Representation for Compositional Reinforcement Learning
Sungbin Mun, Younghwan Lee, Cheolhui MIn, Mineui Hong, Young Min Kim
Abstract
Vision-based reinforcement learning can benefit from object-centric scene representation, which factorizes the visual observation into individual objects and their attributes, such as color, shape, size, and position. While such object-centric representations can extract components that generalize well for various multi-object manipulation tasks, they are prone to issues with occlusions and 3D ambiguity of object properties due to their reliance on single-view 2D image features. Furthermore, the entanglement between object configurations and camera poses complicates the object-centric disentanglement in 3D, leading to poor 3D reasoning by the agent in vision-based reinforcement learning applications. To address the lack of 3D awareness and the object-camera entanglement problem, we propose an enhanced 3D object-centric representation that utilizes multi-view 3D features and enforces more explicit 3D-aware disentanglement. The enhancement is based on the integration of the recent success of multi-view Transformer and the prototypical representation learning among the object-centric representations. The representation, therefore, can stably identify proxies of 3D positions of individual objects along with their semantic and physical properties, exhibiting excellent interpretability and controllability. Then, our proposed block transformer policy effectively performs novel tasks by assembling desired properties adaptive to the new goal states, even when provided with unseen viewpoints at test time. We demonstrate that our 3D-aware block representation is scalable to compose diverse novel scenes and enjoys superior performance in out-of-distribution tasks with multi-object manipulations under both seen and unseen viewpoints compared to existing methods.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 2129231b-ad7a-4453-8ebd-0a03e88ae127Builds on20
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou et al.ICCV 2021 · 8,921 citations
- PaLM-E: An Embodied Multimodal Language ModelDanny Driess, Fei Xia, Mehdi S. M. Sajjadi, Corey Lynch et al.ICML 2023 · 2,601 citations
- Object-Centric Learning with Slot AttentionFrancesco Locatello, Dirk Weissenborn, Thomas Unterthiner, Aravindh Mahendran et al.NeurIPS 2020 · 1,275 citations
- Unsupervised Discovery of Object Radiance FieldsHong-Xing Yu, Leonidas J. Guibas, Jiajun WuICLR 2022 · 132 citations
- Object Scene Representation TransformerMehdi S. M. Sajjadi, Daniel Duckworth, Aravindh Mahendran, Sjoerd van Steenkiste et al.NeurIPS 2022 · 124 citations
Related papers
- An Investigation into Pre-Training Object-Centric Representations for Reinforcement LearningJaesik Yoon, Yi-Fu Wu, Heechul Bae, Sungjin AhnICML 2023 · 59 citations
- Learning Dynamic Attribute-factored World Models for Efficient Multi-object Reinforcement LearningFan Feng, Sara MagliacaneNeurIPS 2023 · 17 citations
- Entity-Centric Reinforcement Learning for Object Manipulation from PixelsDan Haramati, Tal Daniel, Aviv TamarICLR 2024 · 31 citations
- 3D-MVP: 3D Multiview Pretraining for ManipulationShengyi Qian, Kaichun Mo, Valts Blukis, David F. Fouhey et al.CVPR 2025
- Learning to See and Act: Task-Aware Virtual View Exploration for Robotic ManipulationYongjie Bai, Zhouxia Wang, Yang Liu, Kaijun Luo et al.CVPR 2026 · 6 citations
