Lune

NeurIPS2025Top-tier venue

Fine-Grained Preference Optimization Improves Spatial Reasoning in VLMs

Yifan Shen, Yuanzhe Liu, Jingyuan Zhu, Xu Cao, Xiaofeng Zhang, Yixiao He, Wenming Ye, James M. Rehg, Ismini Lourentzou

2025Year
41Citations
14Top-tier citations

Abstract

Current Vision-Language Models (VLMs) struggle with fine-grained spatial reasoning, particularly when multi-step logic and precise spatial alignment are required. In this work, we introduce SpatialReasoner-R1, a vision-language reasoning model designed to address these limitations. To construct high-quality supervision for spatial reasoning, we design a Multi-Model Monte Carlo Tree Search (M3CTS) method that generates diverse, logically consistent Long Chain-of-Thought (Long-CoT) reasoning trajectories. In addition, we propose a fine-grained Direct Preference Optimization (fDPO) method that introduces segment-specific preference 39th Conference on Neural Information Processing Systems (NeurIPS 2025). granularity for descriptive grounding and logical reasoning, guided by a spatial reward mechanism that evaluates candidate responses based on visual consistency, spatial grounding, and logical coherence. Experimental results demonstrate that fDPO achieves relative performance gains of 4.1% and 9.0% over standard DPO on spatial qualitative and quantitative tasks, respectively. SpatialReasoner-R1, trained with fDPO, sets a new SoTA on SPATIALRGPT-BENCH, outperforming the strongest baseline by 9.4% in average accuracy, while maintaining competitive performance on general vision-language tasks.

2 Related Work Vision Language Models and Spatial Reasoning. Recent advances in VLMs have significantly enhanced the ability of multimodal models to understand and generate descriptive text grounded in visual contexts [31,39,40,49,66,95]. Models such as Flamingo [1], BLIP-2 [32], and Qwen-VL [39] use high-capacity vision encoders [53] paired with LLMs [5,64] to achieve state-of-the-art performance in various multimodal tasks, such as visual question answering, image captioning, and instruction following [2,15,37,65,76,100]. Current trends involve scaling models to improve general understanding [12,25,62] and using large-scale instruction tuning datasets [40,56,93]. Both proprietary [21,26,25] and open-source VLMs [12,17,89] have shown impressive results.

While VLMs show promise in visual understanding, accurately perceiving and reasoning about spatial arrangements remains a challenge [13]. Recent efforts to improve spatial understanding include fine-tuning VLMs on spatial VQA datasets [7,8,13,41,75,28,55], and zero-shot frameworks that leverage external 3D foundation models for geometric priors [44]. Region-aware models have also been proposed for better grounding and finer spatial queries [24,87,91]. These advances extend to scenarios such as video understanding [81] and 3D generation [46,50]. To track progress, specialized benchmarks like Q-Spatial Bench [36], SpatialRGPT-Bench [13], VSI-Bench [81], and 3DSRBench [45] have been introduced to assess spatial skills. However, current models still struggle with complex, multi-step spatial reasoning. SpatialReasoner-R1 addresses this gap by introducing fine-grained preference optimization and multi-level reward mechanisms.

Preference-based learning methods, particularly DPO [54], have become standard techniques for aligning models with human intentions. These methods bypass the need for explicit reward model training and have often demonstrated strong performance compared to earlier Reinforcement Learning with Human Feedback (RLHF) approaches [3,19,48,98]. In the multimodal domain, DPO and its variants have been adapted to address specific challenges such as reducing hallucinations and improving visual grounding [70,78,88]. The adaptability of DPO is further highlighted by its recent application in aligning generative models beyond language, such as text-to-image diffusion models [22,33,67,82,90]. Adaptation methods often involve constructing preference pairs based on human corrections, AI feedback, or contrasting inputs to guide the model towards desired behaviors [11,14,18,20,61,68,74,77,79,85].

Standard DPO methods treat the reasoning process as a single structure. To address this, preference granularity in DPO has been explored at the token [38, 57, 94, 97, 99], step [29, 96], sentence [51,54,58], and turn [59, 60, 80] levels. While effective in certain domains, these approaches overlook the semantic roles of different segments in LongCoT, where descriptive grounding and logical reasoning require distinct optimization. In contrast, our proposed fDPO introduces functional-level preference granularity.

Multi-LLM Guided Reasoning Recent work has explored leveraging multiple LLMs to collaboratively solve complex reasoning tasks, often integrated with Monte Carlo Tree Search (MCTS). Methods such as MoA [69], MoSA [84], AlphaLLM-CPL [71], and LE-MCTS [52] enhance multiagent text-based reasoning using ensemble methods and stepwise search. CoMCTS (Mulberry) [86]

extends multi-LLM MCTS to multimodal reasoning, primarily targeting collaborative reflection and error correction. In contrast, our method, M3CTS, addresses the challenge of spatial reasoning in VLMs, introducing fine

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

lune papers fulltext 698edd62-d895-470a-a1ed-4c7389f87bce

Cited by top-tier papers14

Ask how each one uses it

Builds on59

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines