Ascending the Infinite Ladder: Benchmarking Spatial Deformation Reasoning in Vision-Language Models
Jiahuan Zhang, Shunwen Bai, Tianheng Wang, Kaiwen Guo, Zijia Song, Hanqing Wu, Guozheng Rao, Kai Han, Kaicheng Yu
摘要
Under review. (given the final state, determine the operations). We adopt a ladder competition format, using the number of deformation steps as the level classification criterion, with the goal of exploring the boundaries of the model's deformation reasoning capabilities. Interestingly, the benchmarking results reveal that almost no model demonstrates plausible spatial deformation reasoning abilities. Furthermore, even after applying targeted training and mainstream reasoning enhancement methods, the models are still unable to perform well on 3D spatial deformation reasoning. Spatial Deformation Reasoning In this study, spatial deformation reasoning refers to the model's ability to understand, predict, and execute complex deformations of an object's shape. We focus on implementing this reasoning in VLMs, especially when there is no prior knowledge, and the model must learn and perform deformations through observation or manipulation. Unlike visual spatial intelligence in the VSI-Bench [56] , which emphasizes localization, relational understanding, and spatial perception, our spatial deformation centers on dynamic shape transformations, particularly in multi-step processes that change an object's state. We categorize the core capabilities of spatial deformation as follows: Spatial Recognition. The ability to comprehend the initial shape of an object and accurately identify the areas that require deformation. Abstraction of Operational Law Principles. The ability to grasp the principles behind deformation operations and abstract them into deformation laws, crucial for accurate reasoning. Stable Reasoning Execution. The ability to gradually execute the inferred reasoning rules, maintain stability throughout the process, and derive the final state.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper18
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 被引用 11,349 次
- Tree of Thoughts: Deliberate Problem Solving with Large Language ModelsShunyu Yao, Dian Yu, Jeffrey Zhao, Izhak Shafran 等NeurIPS 2023 · 被引用 5,068 次
- Graph of Thoughts: Solving Elaborate Problems with Large Language ModelsMaciej Besta, Nils Blach, Ales Kubicek, Robert Gerstenberger 等AAAI 2024 · 被引用 1,292 次
- AlphaZero-Like Tree-Search can Guide Large Language Model Decoding and TrainingZiyu Wan, Xidong Feng, Muning Wen, Stephen Marcus McAleer 等ICML 2024 · 被引用 325 次
相关 Paper
- 3DSRBENCH: A Comprehensive 3D Spatial Reasoning BenchmarkWufei Ma, Haoyu Chen, Guofeng Zhang, Yu-Cheng Chou 等ICCV 2025 · 被引用 15 次
- Thinking in Structures: Evaluating Spatial Intelligence in Constraint-Governed SpacesChen Yang, Guanxin Lin, Youquan He, Peiyao Chen 等ICML 2026
- SpatialLadder: Progressive Training for Spatial Reasoning in Vision-Language ModelsHongxing Li, Dingming Li, Zixuan Wang, Yuchen Yan 等ICLR 2026 · 被引用 75 次
- Spatial-DISE: A Unified Benchmark for Evaluating Spatial Reasoning in Vision-Language ModelsXinmiao Huang, Qisong He, Zhenglin Huang, Boxuan Wang 等ICLR 2026 · 被引用 8 次
- SpatialPIN: Enhancing Spatial Reasoning Capabilities of Vision-Language Models through Prompting and Interacting 3D PriorsChenyang Ma, Kai Lu, Ta Ying Cheng, Niki Trigoni 等NeurIPS 2024 · 被引用 82 次
