ESPRIT: Explaining Solutions to Physical Reasoning Tasks
Nazneen Fatema Rajani, Rui Zhang, Yi Chern Tan, Stephan Zheng, Jeremy Weiss, Aadit Vyas, Abhijit Gupta, Caiming Xiong, Richard Socher, Dragomir R. Radev
Abstract
Neural networks lack the ability to reason about qualitative physics and so cannot generalize to scenarios and tasks unseen during training. We propose ESPRIT, a framework for commonsense reasoning about qualitative physics in natural language that generates interpretable descriptions of physical events. We use a two-step approach of first identifying the pivotal physical events in an environment and then generating natural language descriptions of those events using a data-to-text approach. Our framework learns to generate explanations of how the physical simulation will causally evolve so that an agent or a human can easily reason about a solution using those interpretable descriptions. Human evaluations indicate that ESPRIT produces crucial fine-grained details and has high coverage of physical concepts compared to even human annotations. Dataset, code and documentation are available at https://github.com/ salesforce/esprit .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 357fbddb-d33e-49b1-a09d-43226c9e7991Cited by top-tier papers6
- ComPhy: Compositional Physical Reasoning of Objects and Events from VideosZhenfang Chen, Kexin Yi, Yunzhu Li, Mingyu Ding et al.ICLR 2022 · 67 citations
- ContPhy: Continuum Physical Concept Learning and Reasoning from VideosZhicheng Zheng, Xin Yan, Zhenfang Chen, Jingzhou Wang et al.ICML 2024 · 22 citations
- PhysInOne: Visual Physics Learning and Reasoning in One SuiteSiyuan Zhou, Hejun Wang, Hu Cheng, Jinxi Li et al.CVPR 2026 · 9 citations
- CRIPP-VQA: Counterfactual Reasoning about Implicit Physical Properties via Video Question AnsweringMaitreya Patel, Tejas Gokhale, Chitta Baral, Yezhou YangEMNLP 2022 · 5 citations
- Dynamics: Language-Based Representation for Inferring Rigid-Body Dynamics From VideosChia-Hsiang Kao, Cong Phuoc Huynh, Chien-Yi Wang, Noranart Vesdapunt et al.CVPR 2026 · 2 citations
Builds on6
- PIQA: Reasoning about Physical Commonsense in Natural LanguageYonatan Bisk, Rowan Zellers, Ronan Le Bras, Jianfeng Gao et al.AAAI 2020 · 2,916 citations
- CLEVRER: Collision Events for Video Representation and ReasoningKexin Yi, Chuang Gan, Yunzhu Li, Pushmeet Kohli et al.ICLR 2020 · 584 citations
- Logical Natural Language Generation from Open-Domain TablesWenhu Chen, Jianshu Chen, Yu Su, Zhiyu Chen et al.ACL 2020 · 116 citations
- Bridging the Structural Gap Between Encoding and Decoding for Data-To-Text GenerationChao Zhao, Marilyn A. Walker, Snigdha ChaturvediACL 2020 · 82 citations
- RTFM: Generalising to New Environment Dynamics via ReadingVictor Zhong, Tim Rocktäschel, Edward GrefenstetteICLR 2020 · 44 citations
Related papers
- A Bayesian-Symbolic Approach to Reasoning and Learning in Intuitive PhysicsKai Xu, Akash Srivastava, Dan Gutfreund, Felix Sosa et al.NeurIPS 2021 · 29 citations
- CueTip: An Interactive and Explainable Physics-aware Pool AssistantSean Memery, Kevin Denamganaï, Jiaxin Zhang, Zehai Tu et al.SIGGRAPH 2025
- Chain of Event-Centric Causal Thought for Physically Plausible Video GenerationZixuan Wang, Yixin Hu, Haolan Wang, Feng Chen et al.CVPR 2026 · 8 citations
- SIMPACT: Simulation-Enabled Action Planning using Vision-Language ModelsHaowen Liu, Shaoxiong Yao, Haonan Chen, Jiawei Gao et al.CVPR 2026 · 8 citations
- Pixel2Phys: Distilling Governing Laws from Visual DynamicsRuikun Li, Jun Yao, Yingfan Hua, Shixiang Tang et al.CVPR 2026 · 2 citations
