ProReason: Multi-Modal Proactive Reasoning with Decoupled Eyesight and Wisdom
Jingqi Zhou, Sheng Wang, Jingwei Dong, Kai Liu, Lei Li, Jiahui Gao, Jiyue Jiang, Lingpeng Kong, Chuan Wu
Abstract
Large vision-language models (LVLMs) have witnessed significant progress on visual understanding tasks. However, they often prioritize language knowledge over image information on visual reasoning tasks, incurring performance degradation. To tackle this issue, we first identify the drawbacks of existing solutions (i.e., limited multi-modal reasoning capacities, and insufficient and irrelevant visual descriptions). We then decompose visual reasoning process into two stages: proactive visual perception (i.e., eyesight) and textual reasoning (i.e., wisdom), and introduce a novel visual reasoning framework named PROREASON. This framework features decoupled vision-reasoning capabilities and multi-run proactive perception. Briefly, given a multi-modal question, PRORE-ASON iterates proactive information collection and reasoning until the answer can be concluded with necessary and sufficient visual descriptions. Notably, the disassociation of capabilities allows seamless integration of existing large language models (LLMs) to compensate for the reasoning deficits of LVLMs. Our extensive experiments demonstrate that PROREASON outperforms existing multi-step reasoning frameworks on various benchmarks for both open-source and closed-source models, with the average performance gain reaching 13.2%. Besides, the integration of LLMs allows PROREASON to produce high-quality visual reasoning data, which empowers PRORE-ASON-distilled models (i.e., ProReason-VL and ProReason-Q3) to achieve superior performance in downstream tasks. Our insights into existing solutions and the decoupled perspective for feasible integration of LLMs illuminate future research on visual reasoning techniques, especially LLM-assisted ones. The code is available at https://github.com/ lian-tian-mo-zun/Pro_Reason .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 7dd521f4-d894-4173-9334-6b2d980d483fCited by top-tier papers4
- Perception-R1: Advancing Multimodal Reasoning Capabilities of MLLMs via Visual Perception RewardTong Xiao, Xin Xu, Zhenya Huang, Hongyu Gao et al.ICLR 2026 · 33 citations
- Look Again, Think Slowly: Enhancing Visual Reflection in Vision-Language ModelsPu Jian, Junhong Wu, Wei Sun, Chen Wang et al.EMNLP 2025 · 2 citations
- A Progressive Visual-Logic-Aligned Framework for Ride-Hailing AdjudicationWeiming Wu, Zi-Jian Cheng, Jie Meng, Peng Zhen et al.KDD 2026
- LC-ERD: Mining Latent Logic for Self-Evolving Reasoning via Consistency-Regulated Reward DecompositionYanyu Chen, Jiyue Jiang, Dianzhi Yu, Zheng Wu et al.KDD 2026
Builds on10
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Tree of Thoughts: Deliberate Problem Solving with Large Language ModelsShunyu Yao, Dian Yu, Jeffrey Zhao, Izhak Shafran et al.NeurIPS 2023 · 5,068 citations
- MMMU: A Massive Multi-Discipline Multimodal Understanding and Reasoning Benchmark for Expert AGIXiang Yue, Yuansheng Ni, Tianyu Zheng, Kai Zhang et al.CVPR 2024 · 213 citations
- Character-LLM: A Trainable Agent for Role-PlayingYunfan Shao, Linyang Li, Junqi Dai, Xipeng QiuEMNLP 2023 · 97 citations
Related papers
- Integrating Visual Interpretation and Linguistic Reasoning for Geometric Problem SolvingZixian Guo, Ming Liu, Qilong Wang, Zhilong Ji et al.ICCV 2025 · 1 citation
- From the Least to the Most: Building a Plug-and-Play Visual Reasoner via Data SynthesisChuanqi Cheng, Jian Guan, Wei Wu, Rui YanEMNLP 2024 · 1 citation
- Enhancing Advanced Visual Reasoning Ability of Large Language ModelsZhiyuan Li, Dongnan Liu, Chaoyi Zhang, Heng Wang et al.EMNLP 2024 · 10 citations
- VisionReasoner: Unified Reasoning-Integrated Visual Perception via Reinforcement LearningYuqi Liu, Tianyuan Qu, Zhisheng Zhong, Bohao Peng et al.ICLR 2026 · 15 citations
- Cantor: Inspiring Multimodal Chain-of-Thought of MLLMTimin Gao, Peixian Chen, Mengdan Zhang, Chaoyou Fu et al.ACM MM 2024 · 20 citations
