Intrinsic Physical Concepts Discovery with Object-Centric Predictive Models
Qu Tang, Xiangyu Zhu, Zhen Lei, Zhaoxiang Zhang
Abstract
The ability to discover abstract physical concepts and understand how they work in the world through observing lies at the core of human intelligence. The acquisition of this ability is based on compositionally perceiving the environment in terms of objects and relations in an unsupervised manner. Recent approaches learn object-centric representations and capture visually observable concepts of objects, e.g., shape, size, and location. In this paper, we take a step forward and try to discover and represent intrinsic physical concepts such as mass and charge. We introduce the PHYsical Concepts Inference NEtwork (PHYCINE), a system that infers physical concepts in different abstract levels without supervision. The key insights underlining PHYCINE are two-fold, commonsense knowledge emerges with prediction, and physical concepts of different abstract levels should be reasoned in a bottom-up fashion. Empirical evaluation demonstrates that variables inferred by our system work in accordance with the properties of the corresponding physical concepts. We also show that object representations containing the discovered physical concepts variables could help achieve better performance in causal reasoning tasks, i.e., ComPhy.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 4d9a5772-7080-4d80-ab4a-e6cd9ae34f15Cited by top-tier papers2
- Reasoning-Enhanced Object-Centric Learning for VideosJian Li, Pu Ren, Yang Liu, Hao SunKDD 2025 · 7 citations
- Ock: Unsupervised Dynamic Video Prediction With Object-Centric KinematicsYeon-Ji Song, Jaein Kim, Suhyung Choi, Jin-Hwa Kim et al.ICCV 2025 · 4 citations
Builds on11
- Object-Centric Learning with Slot AttentionFrancesco Locatello, Dirk Weissenborn, Thomas Unterthiner, Aravindh Mahendran et al.NeurIPS 2020 · 1,275 citations
- CLEVRER: Collision Events for Video Representation and ReasoningKexin Yi, Chuang Gan, Yunzhu Li, Pushmeet Kohli et al.ICLR 2020 · 584 citations
- Contrastive Learning of Structured World ModelsThomas N. Kipf, Elise van der Pol, Max WellingICLR 2020 · 322 citations
- Simple Unsupervised Object-Centric Learning for Complex and Naturalistic VideosGautam Singh, Yi-Fu Wu, Sungjin AhnNeurIPS 2022 · 182 citations
- Dynamic Visual Reasoning by Learning Differentiable Physics Models from Video and LanguageMingyu Ding, Zhenfang Chen, Tao Du, Ping Luo et al.NeurIPS 2021 · 90 citations
Related papers
- Hierarchical Relational InferenceAleksandar Stanic, Sjoerd van Steenkiste, Jürgen SchmidhuberAAAI 2021 · 17 citations
- ComPhy: Compositional Physical Reasoning of Objects and Events from VideosZhenfang Chen, Kexin Yi, Yunzhu Li, Mingyu Ding et al.ICLR 2022 · 67 citations
- Visual Grounding of Learned Physical ModelsYunzhu Li, Toru Lin, Kexin Yi, Daniel Bear et al.ICML 2020 · 88 citations
- Unsupervised Discovery of 3D Physical Objects from VideoYilun Du, Kevin A. Smith, Tomer D. Ullman, Joshua B. Tenenbaum et al.ICLR 2021 · 11 citations
- Pixel2Phys: Distilling Governing Laws from Visual DynamicsRuikun Li, Jun Yao, Yingfan Hua, Shixiang Tang et al.CVPR 2026 · 2 citations
