Fine-Grained Preference Optimization Improves Spatial Reasoning in VLMs
Yifan Shen, Yuanzhe Liu, Jingyuan Zhu, Xu Cao, Xiaofeng Zhang, Yixiao He, Wenming Ye, James M. Rehg, Ismini Lourentzou
Abstract
Current Vision-Language Models (VLMs) struggle with fine-grained spatial reasoning, particularly when multi-step logic and precise spatial alignment are required. In this work, we introduce SpatialReasoner-R1, a vision-language reasoning model designed to address these limitations. To construct high-quality supervision for spatial reasoning, we design a Multi-Model Monte Carlo Tree Search (M3CTS) method that generates diverse, logically consistent Long Chain-of-Thought (Long-CoT) reasoning trajectories. In addition, we propose a fine-grained Direct Preference Optimization (fDPO) method that introduces segment-specific preference 39th Conference on Neural Information Processing Systems (NeurIPS 2025). granularity for descriptive grounding and logical reasoning, guided by a spatial reward mechanism that evaluates candidate responses based on visual consistency, spatial grounding, and logical coherence. Experimental results demonstrate that fDPO achieves relative performance gains of 4.1% and 9.0% over standard DPO on spatial qualitative and quantitative tasks, respectively. SpatialReasoner-R1, trained with fDPO, sets a new SoTA on SPATIALRGPT-BENCH, outperforming the strongest baseline by 9.4% in average accuracy, while maintaining competitive performance on general vision-language tasks.
2 Related Work Vision Language Models and Spatial Reasoning. Recent advances in VLMs have significantly enhanced the ability of multimodal models to understand and generate descriptive text grounded in visual contexts [31,39,40,49,66,95]. Models such as Flamingo [1], BLIP-2 [32], and Qwen-VL [39] use high-capacity vision encoders [53] paired with LLMs [5,64] to achieve state-of-the-art performance in various multimodal tasks, such as visual question answering, image captioning, and instruction following [2,15,37,65,76,100]. Current trends involve scaling models to improve general understanding [12,25,62] and using large-scale instruction tuning datasets [40,56,93]. Both proprietary [21,26,25] and open-source VLMs [12,17,89] have shown impressive results.
While VLMs show promise in visual understanding, accurately perceiving and reasoning about spatial arrangements remains a challenge [13]. Recent efforts to improve spatial understanding include fine-tuning VLMs on spatial VQA datasets [7,8,13,41,75,28,55], and zero-shot frameworks that leverage external 3D foundation models for geometric priors [44]. Region-aware models have also been proposed for better grounding and finer spatial queries [24,87,91]. These advances extend to scenarios such as video understanding [81] and 3D generation [46,50]. To track progress, specialized benchmarks like Q-Spatial Bench [36], SpatialRGPT-Bench [13], VSI-Bench [81], and 3DSRBench [45] have been introduced to assess spatial skills. However, current models still struggle with complex, multi-step spatial reasoning. SpatialReasoner-R1 addresses this gap by introducing fine-grained preference optimization and multi-level reward mechanisms.
Preference-based learning methods, particularly DPO [54], have become standard techniques for aligning models with human intentions. These methods bypass the need for explicit reward model training and have often demonstrated strong performance compared to earlier Reinforcement Learning with Human Feedback (RLHF) approaches [3,19,48,98]. In the multimodal domain, DPO and its variants have been adapted to address specific challenges such as reducing hallucinations and improving visual grounding [70,78,88]. The adaptability of DPO is further highlighted by its recent application in aligning generative models beyond language, such as text-to-image diffusion models [22,33,67,82,90]. Adaptation methods often involve constructing preference pairs based on human corrections, AI feedback, or contrasting inputs to guide the model towards desired behaviors [11,14,18,20,61,68,74,77,79,85].
Standard DPO methods treat the reasoning process as a single structure. To address this, preference granularity in DPO has been explored at the token [38, 57, 94, 97, 99], step [29, 96], sentence [51,54,58], and turn [59, 60, 80] levels. While effective in certain domains, these approaches overlook the semantic roles of different segments in LongCoT, where descriptive grounding and logical reasoning require distinct optimization. In contrast, our proposed fDPO introduces functional-level preference granularity.
Multi-LLM Guided Reasoning Recent work has explored leveraging multiple LLMs to collaboratively solve complex reasoning tasks, often integrated with Monte Carlo Tree Search (MCTS). Methods such as MoA [69], MoSA [84], AlphaLLM-CPL [71], and LE-MCTS [52] enhance multiagent text-based reasoning using ensemble methods and stepwise search. CoMCTS (Mulberry) [86]
extends multi-LLM MCTS to multimodal reasoning, primarily targeting collaborative reflection and error correction. In contrast, our method, M3CTS, addresses the challenge of spatial reasoning in VLMs, introducing fine
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 698edd62-d895-470a-a1ed-4c7389f87bceCited by top-tier papers14
- DriveMoE: Mixture-of-Experts for Vision-Language-Action Model in End-to-End Autonomous DrivingZhenjie Yang, Yilin Chai, Xiaosong Jia, Qifeng Li et al.CVPR 2026 · 108 citations
- PALM: Progress-Aware Policy Learning via Affordance Reasoning for Long-Horizon Robotic ManipulationYuanzhe Liu, Jingyuan Zhu, Yuchen Mo, Gen Li et al.CVPR 2026 · 31 citations
- FusionAgent: A Multimodal Agent with Dynamic Model Selection for Human RecognitionJie Zhu, Xiao Guo, Yiyang Su, Anil K. Jain et al.CVPR 2026 · 7 citations
- ViBES: A Conversational Agent with Behaviorally-Intelligent 3D Virtual BodyJuze Zhang, Changan Chen, Xin Chen, Heng Yu et al.CVPR 2026 · 7 citations
- PCA-Seg: Revisiting Cost Aggregation for Open-Vocabulary Semantic and Part SegmentationJianjian Yin, Tao Chen, Yi Chen, Gensheng Pei et al.CVPR 2026 · 6 citations
Builds on59
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning et al.NeurIPS 2023 · 10,924 citations
Related papers
- SpatialReasoner: Towards Explicit and Generalizable 3D Spatial ReasoningWufei Ma, Yu-Cheng Chou, Qihao Liu, Xingrui Wang et al.NeurIPS 2025 · 77 citations
- Reinforcing Video Reasoning Segmentation to Think Before It SegmentsSitong Gong, Yunzhi Zhuge, Lu Zhang, Jiazuo Yu et al.CVPR 2026 · 16 citations
- STAR-R1: Multi-View Spatial TrAnsformation Reasoning by Reinforcing Multimodal LLMsZongzhao Li, Zongyang Ma, Mingze Li, Songyou Li et al.CVPR 2026
- OmniSpatial: Towards Comprehensive Spatial Reasoning Benchmark for Vision Language ModelsMengdi Jia, Zekun Qi, Shaochen Zhang, Wenyao Zhang et al.ICLR 2026 · 109 citations
- Enhancing Spatial Reasoning Through Visual and Textual ThinkingXun Liang, Xin Guo, Zhongming Jin, Weihang Pan et al.AAAI 2026
