Multi-step Visual Reasoning with Visual Tokens Scaling and Verification
Tianyi Bai, Zengjie Hu, Fupeng Sun, Jiantao Qiu, Yizhen Jiang, Guangxin He, Bohan Zeng, Conghui He, Binhang Yuan, Wentao Zhang
摘要
Multi-modal large language models (MLLMs) have achieved remarkable capabilities by integrating visual perception with language understanding, enabling applications such as image-grounded dialogue, visual question answering, and scientific analysis. However, most MLLMs adopt a static inference paradigm, encoding the entire image into fixed visual tokens upfront, which limits their ability to iteratively refine understanding or adapt to context during inference. This contrasts sharply with human perception, which is dynamic, selective, and feedback-driven. In this work, we introduce a novel framework for inference-time visual token scaling that enables MLLMs to perform iterative, verifier-guided reasoning over visual content. We formulate the problem as a Markov Decision Process, involving a reasoner that proposes visual actions and a verifier-trained via multi-step Direct Preference Optimization (DPO)-that evaluates these actions and determines when reasoning should terminate. To support this, we present a new dataset, VTS, comprising supervised reasoning trajectories (VTS-SFT) and preference-labeled reasoning comparisons (VTS-DPO). Our method significantly outperforms existing approaches across diverse visual reasoning benchmarks, offering not only improved accuracy but also more interpretable and grounded reasoning processes. These results demonstrate the promise of dynamic inference mechanisms for enabling fine-grained, context-aware visual reasoning in next-generation MLLMs. Code and datasets are publicly released at https://vts-v.github.io/.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- VABench: A Comprehensive Benchmark for Audio-Video GenerationDaili Hua, Xizhi Wang, Bohan Zeng, Xinyi Huang 等CVPR 2026 · 被引用 25 次
- Synthesizing Multimodal Geometry Datasets from Scratch and Enabling Visual Alignment via Plotting CodeHaobo Lin, Tianyi Bai, Chen Chen, Jiajun Zhang 等ICML 2026 · 被引用 1 次
它引用的顶会 Paper17
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning 等NeurIPS 2023 · 被引用 10,924 次
- Let's Verify Step by StepHunter Lightman, Vineet Kosaraju, Yuri Burda, Harrison Edwards 等ICLR 2024 · 被引用 3,045 次
- MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual ContextsPan Lu, Hritik Bansal, Tony Xia, Jiacheng Liu 等ICLR 2024 · 被引用 1,472 次
- Are We on the Right Way for Evaluating Large Vision-Language Models?Lin Chen, Jinsong Li, Xiaoyi Dong, Pan Zhang 等NeurIPS 2024 · 被引用 1,029 次
相关 Paper
- Unleashing Perception-Time Scaling to Multimodal Reasoning ModelsYifan Li, Zhenghao Chen, Ziheng Wu, Kun Zhou 等ICLR 2026 · 被引用 1 次
- VideoChat-R1.5: Visual Test-Time Scaling to Reinforce Multimodal Reasoning by Iterative PerceptionZiang Yan, Yinan He, Xinhao Li, Zhengrong Yue 等NeurIPS 2025 · 被引用 70 次
- VisRef: Visual Refocusing while Thinking Improves Test-Time Scaling in Multi-Modal Large Reasoning ModelsSoumya Suvra Ghosal, Youngeun Kim, Zhuowei Li, Ritwick Chaudhry 等CVPR 2026
- Vision-aligned Latent Reasoning for Multi-modal Large Language ModelByungwoo Jeon, Yoonwoo Jeong, Hyunseok Lee, Minsu Cho 等ICML 2026 · 被引用 7 次
- Spotlight on Token Perception for Multimodal Reinforcement LearningSiyuan Huang, Xiaoye Qu, Yafu Li, Yun Luo 等ICLR 2026 · 被引用 45 次
