From the Least to the Most: Building a Plug-and-Play Visual Reasoner via Data Synthesis
Chuanqi Cheng, Jian Guan, Wei Wu, Rui Yan
摘要
We explore multi-step reasoning in visionlanguage models (VLMs). The problem is challenging, as reasoning data consisting of multiple steps of visual and language processing are barely available. To overcome the challenge, we first introduce a least-to-most visual reasoning paradigm, which interleaves steps of decomposing a question into sub-questions and invoking external tools for resolving subquestions. Based on the paradigm, we further propose a novel data synthesis approach that can automatically create questions and multistep reasoning paths for an image in a bottomup manner. Our approach divides the complex synthesis task into a few simple sub-tasks, and (almost entirely) relies on open-sourced models to accomplish the sub-tasks. Therefore, the entire synthesis process is reproducible and cost-efficient, and the synthesized data is quality guaranteed. With the approach, we construct 50k visual reasoning examples. Then, we develop a visual reasoner through supervised fine-tuning, which is capable of generally enhancing the reasoning abilities of a wide range of existing VLMs in a plug-and-play fashion. Extensive experiments indicate that the visual reasoner can consistently and significantly improve four VLMs on four VQA benchmarks. Our code and dataset are available at https:// github.com/steven-ccq/VisualReasoner .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper8
- Reinforcing Spatial Reasoning in Vision-Language Models with Interwoven Thinking and Visual DrawingJunfei Wu, Jian Guan, Kaituo Feng, Qiang Liu 等NeurIPS 2025 · 被引用 153 次
- AMOR: A Recipe for Building Adaptable Modular Knowledge Agents Through Process FeedbackJian Guan, Wei Wu, Zujie Wen, Peng Xu 等NeurIPS 2024 · 被引用 35 次
- Weaving Context Across Images: Improving Vision-Language Models through Focus-Centric Visual ChainsJuntian Zhang, Chuanqi Cheng, Yuhan Liu, Wei Liu 等ACL 2025 · 被引用 13 次
- A Unified Agentic Framework for Evaluating Conditional Image GenerationJifang Wang, Xue Yang, Longyue Wang, Zhenran Xu 等ACL 2025 · 被引用 7 次
- 2D-TPE: Two-Dimensional Positional Encoding Enhances Table Understanding for Large Language ModelsJia-Nan Li, Jian Guan, Wei Wu, Zhengtao Yu 等WWW 2025 · 被引用 4 次
它引用的顶会 Paper20
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Large Language Models are Zero-Shot ReasonersTakeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo 等NeurIPS 2022 · 被引用 8,168 次
- Deformable DETR: Deformable Transformers for End-to-End Object DetectionXizhou Zhu, Weijie Su, Lewei Lu, Bin Li 等ICLR 2021 · 被引用 7,353 次
- Flamingo: a Visual Language Model for Few-Shot LearningJean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech 等NeurIPS 2022 · 被引用 6,707 次
- Toolformer: Language Models Can Teach Themselves to Use ToolsTimo Schick, Jane Dwivedi-Yu, Roberto Dessì, Roberta Raileanu 等NeurIPS 2023 · 被引用 5,989 次
相关 Paper
- ProReason: Multi-Modal Proactive Reasoning with Decoupled Eyesight and WisdomJingqi Zhou, Sheng Wang, Jingwei Dong, Kai Liu 等EMNLP 2025
- Integrating Visual Interpretation and Linguistic Reasoning for Geometric Problem SolvingZixian Guo, Ming Liu, Qilong Wang, Zhilong Ji 等ICCV 2025 · 被引用 1 次
- Through the Lens of Contrast: Self-Improving Visual Reasoning in VLMsZhiyu Pan, Yizheng Wu, Jiashen Hua, Junyi Feng 等ICLR 2026 · 被引用 11 次
- Generalizing from SIMPLE to HARD Visual Reasoning: Can We Mitigate Modality Imbalance in VLMs?Simon Park, Abhishek Panigrahi, Yun Cheng, Dingli Yu 等ICML 2025
- Enhancing Cognition and Explainability of Multimodal Foundation Models with Self-Synthesized DataYucheng Shi, Quanzheng Li, Jin Sun, Xiang Li 等ICLR 2025
