PTR: A Benchmark for Part-based Conceptual, Relational, and Physical Reasoning
Yining Hong, Li Yi, Josh Tenenbaum, Antonio Torralba, Chuang Gan
Abstract
A critical aspect of human visual perception is the ability to parse visual scenes into individual objects and further into object parts, forming part-whole hierarchies. Such composite structures could induce a rich set of semantic concepts and relations, thus playing an important role in the interpretation and organization of visual signals as well as for the generalization of visual perception and reasoning. However, existing visual reasoning benchmarks mostly focus on objects rather than parts. Visual reasoning based on the full part-whole hierarchy is much more challenging than object-centric reasoning due to finer-grained concepts, richer geometry relations, and more complex physics. Therefore, to better serve for part-based conceptual, relational and physical reasoning, we introduce a new large-scale diagnostic visual reasoning dataset named PTR. PTR contains around 70k RGBD synthetic images with ground truth object and part level annotations regarding semantic instance segmentation, color attributes, spatial and geometric relationships, and certain physical properties such as stability. These images are paired with 700k machine-generated questions covering various types of reasoning types, making them a good testbed for visual reasoning models. We examine several state-of-the-art visual reasoning models on this dataset and observe that they still make many surprising mistakes in situations where humans can easily infer the correct answer. We believe this dataset will open up new opportunities for part-based reasoning. PTR dataset and baseline models are publicly available 2 .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 7ceebf4c-76a9-4c72-8c88-84e4845bd5a3Cited by top-tier papers16
- ComPhy: Compositional Physical Reasoning of Objects and Events from VideosZhenfang Chen, Kexin Yi, Yunzhu Li, Mingyu Ding et al.ICLR 2022 · 67 citations
- RISP: Rendering-Invariant State Predictor with Differentiable Simulation and Rendering for Cross-Domain Parameter EstimationPingchuan Ma, Tao Du, Joshua B. Tenenbaum, Wojciech Matusik et al.ICLR 2022 · 36 citations
- Contact Points Discovery for Soft-Body Manipulations with Differentiable PhysicsSizhe Li, Zhiao Huang, Tao Du, Hao Su et al.ICLR 2022 · 30 citations
- 3D-Aware Visual Question Answering about Parts, Poses and OcclusionsXingrui Wang, Wufei Ma, Zhuowan Li, Adam Kortylewski et al.NeurIPS 2023 · 27 citations
- Interpretable part-whole hierarchies and conceptual-semantic relationships in neural networksNicola Garau, Niccolò Bisagno, Zeno Sambugaro, Nicola ConciCVPR 2022 · 22 citations
Builds on15
- Object-Centric Learning with Slot AttentionFrancesco Locatello, Dirk Weissenborn, Thomas Unterthiner, Aravindh Mahendran et al.NeurIPS 2020 · 1,275 citations
- MDETR - Modulated Detection for End-to-End Multi-Modal UnderstandingAishwarya Kamath, Mannat Singh, Yann LeCun, Gabriel Synnaeve et al.ICCV 2021 · 1,114 citations
- CLEVRER: Collision Events for Video Representation and ReasoningKexin Yi, Chuang Gan, Yunzhu Li, Pushmeet Kohli et al.ICLR 2020 · 584 citations
- Language-Conditioned Graph Networks for Relational ReasoningRonghang Hu, Anna Rohrbach, Trevor Darrell, Kate SaenkoICCV 2019 · 183 citations
- Closed Loop Neural-Symbolic Learning via Integrating Neural Perception, Grammar Parsing, and Symbolic ReasoningQing Li, Siyuan Huang, Yining Hong, Yixin Chen et al.ICML 2020 · 93 citations
Related papers
- MMR: A Large-scale Benchmark Dataset for Multi-target and Multi-granularity Reasoning SegmentationDonggon Jang, Yucheol Cho, Suin Lee, Taehyeon Kim et al.ICLR 2025
- VisuLogic: A Benchmark for Evaluating Visual Reasoning in Multi-modal Large Language ModelsWeiye Xu, Jiahao Wang, Weiyun Wang, Zhe Chen et al.ICLR 2026 · 103 citations
- Motion-Grounded Video Reasoning: Understanding and Perceiving Motion at Pixel LevelAndong Deng, Tongjia Chen, Shoubin Yu, Taojiannan Yang et al.CVPR 2025
- A Very Big Video Reasoning SuiteMaijunxian Wang, Ruisi Wang, Juyi Lin, Ran Ji et al.ICML 2026 · 20 citations
- PARSE: Part-Aware Relational Spatial ModelingYinuo Bai, Peijun Xu, Kuixiang Shao, Yuyang Jiao et al.CVPR 2026
