Small Drafts, Big Verdict: Information-Intensive Visual Reasoning via Speculation
Yuhan Liu, Lianhui Qin, Shenji Wan
Abstract
Large Vision-Language Models (VLMs) have achieved remarkable progress in multimodal understanding, yet they struggle when reasoning over informationintensive images that densely interleave textual annotations with fine-grained graphical elements. The main challenges lie in precisely localizing critical cues in dense layouts and multi-hop reasoning to integrate dispersed evidence. We propose Speculative Verdict (SV), a training-free framework inspired by speculative decoding that combines multiple lightweight draft experts with a large verdict model. In the draft stage, small VLMs act as draft experts to generate reasoning paths that provide diverse localization candidates; in the verdict stage, a strong VLM synthesizes these paths to produce the final answer, minimizing computational cost while recovering correct answers. To further improve efficiency and accuracy, SV introduces a consensus expert selection mechanism that forwards only high-agreement reasoning paths to the verdict. Empirically, SV achieves consistent gains on challenging information-intensive and high-resolution visual question answering benchmarks, including Infograph-icVQA, ChartMuseum, ChartQAPro, and HR-Bench 4K. By synthesizing correct insights from multiple partially accurate reasoning paths, SV achieves both error correction and cost-efficiency compared to large proprietary models or training pipelines. Code is available at https://github.com/Tinaliu0123/ speculative-verdict .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 8f705ba8-91c9-4e16-904f-0760508fc1b7Builds on19
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- Fast Inference from Transformers via Speculative DecodingYaniv Leviathan, Matan Kalman, Yossi MatiasICML 2023 · 1,472 citations
- Vision-R1: Incentivizing Reasoning Capability in Multimodal Large Language ModelsWenxuan Huang, Bohan Jia, Shaosheng Cao, Zheyu Ye et al.ICLR 2026 · 670 citations
- Visual Sketchpad: Sketching as a Visual Chain of Thought for Multimodal Language ModelsYushi Hu, Weijia Shi, Xingyu Fu, Dan Roth et al.NeurIPS 2024 · 373 citations
- GRIT: Teaching MLLMs to Think with ImagesYue Fan, Xuehai He, Diji Yang, Kaizhi Zheng et al.NeurIPS 2025 · 132 citations
Related papers
- ViSpec: Accelerating Vision-Language Models with Vision-Aware Speculative DecodingJialiang Kang, Han Shu, Wenshuo Li, Yingjie Zhai et al.NeurIPS 2025 · 24 citations
- FastV-RAG: Towards Fast and Fine-Grained Video QA with Retrieval-Augmented GenerationGen Li, Peiyu LiuACL 2026 · 1 citation
- See It, Say It, Sorted: An Iterative Training-Free Framework for Visually-Grounded Multimodal Reasoning in LVLMsYongchang Zhang, Xianzheng Ma, Tianyi Liu, Guangquan Zhou et al.CVPR 2026 · 2 citations
- StaR-KVQA: Structured Reasoning Traces for Implicit-Knowledge Visual Question AnsweringZhihao Wen, Wenkang Wei, Yuan Fang, Xingtong Yu et al.CVPR 2026 · 1 citation
- See the Forest for the Trees: Loosely Speculative Decoding via Visual-Semantic Guidance for Efficient Inference of Video LLMsYicheng Ji, Jun Zhang, Jinpeng Chen, Cong Wang et al.ACL 2026 · 4 citations
