Visual-O1: Understanding Ambiguous Instructions via Multi-modal Multi-turn Chain-of-thoughts Reasoning
Minheng Ni, Yutao Fan, Lei Zhang, Wangmeng Zuo
摘要
As large-scale models evolve, language instructions are increasingly utilized in multi-modal tasks. Due to human language habits, these instructions often contain ambiguities in real-world scenarios, necessitating the integration of visual context or common sense for accurate interpretation. However, even highly intelligent large models exhibit observable performance limitations on ambiguous instructions, where weak reasoning abilities of disambiguation can lead to catastrophic errors. To address this issue, this paper proposes VISUAL-O1, a multimodal multi-turn chain-of-thought reasoning framework. It simulates human multi-modal multi-turn reasoning, providing instantial experience for highly intelligent models or empirical experience for generally intelligent models to understand ambiguous instructions. Unlike traditional methods that require models to possess high intelligence to understand long texts or perform lengthy complex reasoning, our framework does not notably increase computational overhead and is more general and effective, even for generally intelligent models. Experiments show that our method not only enhances the performance of models of different intelligence levels on ambiguous instructions but also improves their performance on general datasets. Our work highlights the potential of artificial intelligence to work like humans in real-world scenarios with uncertainty and ambiguity. We release our data and code at https://github.com/kodenii/Visual-O1.
Recently, chain-of-thoughts (CoT) reasoning has greatly enhanced the understanding and analytical capabilities of high-intelligent large models, such as GPT-4O. However, its application to scenar-
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- VLM-R³: Region Recognition, Reasoning, and Refinement for Enhanced Multimodal Chain-of-ThoughtChaoya Jiang, Yongrui Heng, Wei Ye, Haiyang Xu 等NeurIPS 2025 · 被引用 56 次
- Teaching Vision-Language Models to Ask: Resolving Ambiguity in Visual QuestionsPu Jian, Donglei Yu, Wen Yang, Shuo Ren 等ACL 2025
- Argus: Vision-Centric Reasoning with Grounded Chain-of-ThoughtYunze Man, De-An Huang, Guilin Liu, Shiwei Sheng 等CVPR 2025
- CMMCoT: Enhancing Complex Multi-Image Comprehension via Multi-Modal Chain-of-Thought and Memory AugmentationGuanghao Zhang, Tao Zhong, Yan Xia, Mushui Liu 等AAAI 2026
它引用的顶会 Paper15
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- Tree of Thoughts: Deliberate Problem Solving with Large Language ModelsShunyu Yao, Dian Yu, Jeffrey Zhao, Izhak Shafran 等NeurIPS 2023 · 被引用 5,068 次
- InstructBLIP: Towards General-purpose Vision-Language Models with Instruction TuningWenliang Dai, Junnan Li, Dongxu Li, Anthony Meng Huat Tiong 等NeurIPS 2023 · 被引用 4,013 次
- Unified Vision-Language Pre-Training for Image Captioning and VQALuowei Zhou, Hamid Palangi, Lei Zhang, Houdong Hu 等AAAI 2020 · 被引用 1,047 次
相关 Paper
- Uni-CoT: Towards Unified Chain-of-Thought Reasoning Across Text and VisionLuozheng Qin, Jia Gong, Yuqing Sun, Tianjiao Li 等ICLR 2026 · 被引用 55 次
- DDCoT: Duty-Distinct Chain-of-Thought Prompting for Multimodal Reasoning in Language ModelsGe Zheng, Bin Yang, Jiajin Tang, Hong-Yu Zhou 等NeurIPS 2023 · 被引用 252 次
- Look Light, Think Heavy: What Multimodal Chain-of-Thought Reasoning Can and Cannot DoZhuoran Jin, Kejian Zhu, Hongbang Yuan, Yupu Hao 等ACL 2026
- Fine-Tuning Large Vision-Language Models as Decision-Making Agents via Reinforcement LearningSimon Zhai, Hao Bai, Zipeng Lin, Jiayi Pan 等NeurIPS 2024 · 被引用 214 次
- VLR-Driver: Large Vision-Language-Reasoning Models for Embodied Autonomous DrivingFanjie Kong, Yitong Li, Weihuang Chen, Chen Min 等ICCV 2025 · 被引用 2 次
