On the Learning Mechanisms in Physical Reasoning
Shiqian Li, Kewen Wu, Chi Zhang, Yixin Zhu
Abstract
Is dynamics prediction indispensable for physical reasoning? If so, what kind of roles do the dynamics prediction modules play during the physical reasoning process? Most studies focus on designing dynamics prediction networks and treating physical reasoning as a downstream task without investigating the questions above, taking for granted that the designed dynamics prediction would undoubtedly help the reasoning process. In this work, we take a closer look at this assumption, exploring this fundamental hypothesis by comparing two learning mechanisms: Learning from Dynamics (LfD) and Learning from Intuition (LfI). In the first experiment, we directly examine and compare these two mechanisms. Results show a surprising finding: Simple LfI is better than or on par with state-of-the-art LfD. This observation leads to the second experiment with Ground-truth Dynamics (GD), the ideal case of LfD wherein dynamics are obtained directly from a simulator. Results show that dynamics, if directly given instead of approximated, would achieve much higher performance than LfI alone on physical reasoning; this essentially serves as the performance upper bound. Yet practically, LfD mechanism can only predict Approximate Dynamics (AD) using dynamics learning modules that mimic the physical laws, making the following downstream physical reasoning modules degenerate into the LfI paradigm; see the third experiment. We note that this issue is hard to mitigate, as dynamics prediction errors inevitably accumulate in the long horizon. Finally, in the fourth experiment, we note that LfI, the extremely simpler strategy when done right, is more effective in learning to solve physical reasoning problems. Taken together, the results on the challenging benchmark of PHYRE [3] show that LfI is, if not better, as good as LfD with bells and whistles for dynamics prediction. However, the potential improvement from LfD, though challenging, remains lucrative.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers12
- MEWL: Few-shot multimodal word learning with referential uncertaintyGuangyuan Jiang, Manjie Xu, Shiji Xin, Wei Liang et al.ICML 2023 · 29 citations
- QUANTIPHY: A Quantitative Benchmark Evaluating Physical Reasoning Abilities of Vision-Language ModelsLi Puyin, Tiange Xiang, Ella Mao, Shirley Wei et al.CVPR 2026 · 23 citations
- ContPhy: Continuum Physical Concept Learning and Reasoning from VideosZhicheng Zheng, Xin Yan, Zhenfang Chen, Jingzhou Wang et al.ICML 2024 · 22 citations
- I-PHYRE: Interactive Physical ReasoningShiqian Li, Kewen Wu, Chi Zhang, Yixin ZhuICLR 2024 · 16 citations
- Neural-Symbolic Recursive Machine for Systematic GeneralizationQing Li, Yixin Zhu, Yitao Liang, Ying Nian Wu et al.ICLR 2024 · 15 citations
Builds on16
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- BEiT: BERT Pre-Training of Image TransformersHangbo Bao, Li Dong, Songhao Piao, Furu WeiICLR 2022 · 3,632 citations
- ViViT: A Video Vision TransformerAnurag Arnab, Mostafa Dehghani, Georg Heigold, Chen Sun et al.ICCV 2021 · 2,947 citations
- Is Space-Time Attention All You Need for Video Understanding?Gedas Bertasius, Heng Wang, Lorenzo TorresaniICML 2021 · 2,927 citations
Related papers
- Causal-PIK: Causality-based Physical Reasoning with a Physics-Informed KernelCarlota Parés-Morlans, Michelle Yi, Claire Chen, Sarah A. Wu et al.ICML 2025
- CoPhy: Counterfactual Learning of Physical DynamicsFabien Baradel, Natalia Neverova, Julien Mille, Greg Mori et al.ICLR 2020 · 105 citations
- Beyond Static Vision: Scene Dynamic Field Unlocks Intuitive Physics Understanding in Multi-modal Large Language ModelsNanxi Li, Xiang Wang, Yuanjie Chen, Haode Zhang et al.ICLR 2026
- Dynamic Visual Reasoning by Learning Differentiable Physics Models from Video and LanguageMingyu Ding, Zhenfang Chen, Tao Du, Ping Luo et al.NeurIPS 2021 · 90 citations
- Neural Force Field: Few-shot Learning of Generalized Physical ReasoningShiqian Li, Ruihong Shen, Yaoyu Tao, Chi Zhang et al.ICLR 2026 · 1 citation
