Chain-of-region: Visual Language Models Need Details for Diagram Analysis
Xue Li, Yiyou Sun, Wei Cheng, Yinglun Zhu, Haifeng Chen
Abstract
Visual Language Models (VLMs) like GPT-4V have broadened the scope of LLM applications, yet they face significant challenges in accurately processing visual details, particularly in scientific diagrams. This paper explores the necessity of meticulous visual detail collection and region decomposition for enhancing the performance of VLMs in scientific diagram analysis. We propose a novel approach that combines traditional computer vision techniques with VLMs to systematically decompose diagrams into discernible visual elements and aggregate essential metadata. Our method employs techniques in OpenCV library to identify and label regions, followed by a refinement process using shape detection and region merging algorithms, which are particularly suited to the structured nature of scientific diagrams. This strategy not only improves the granularity and accuracy of visual information processing but also extends the capabilities of VLMs beyond their current limitations. We validate our approach through a series of experiments that demonstrate enhanced performance in diagram analysis tasks, setting a new standard for integrating visual and language processing in a multimodal context.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers1
Ask how each one uses itBuilds on21
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao et al.ICCV 2023 · 13,211 citations
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
- Large Language Models are Zero-Shot ReasonersTakeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo et al.NeurIPS 2022 · 8,168 citations
Related papers
- Primitive Vision: Improving Diagram Understanding in MLLMsShan Zhang, Aotian Chen, Yanpeng Sun, Jindong Gu et al.ICML 2025
- Multimodal ArXiv: A Dataset for Improving Scientific Comprehension of Large Vision-Language ModelsLei Li, Yuqi Wang, Runxin Xu, Peiyi Wang et al.ACL 2024 · 16 citations
- mPLUG-PaperOwl: Scientific Diagram Analysis with the Multimodal Large Language ModelAnwen Hu, Yaya Shi, Haiyang Xu, Jiabo Ye et al.ACM MM 2024 · 15 citations
- Why are Visually-Grounded Language Models Bad at Image Classification?Yuhui Zhang, Alyssa Unell, Xiaohan Wang, Dhruba Ghosh et al.NeurIPS 2024 · 128 citations
- Do Vision-Language Models Really Understand Visual Language?Yifan Hou, Buse Giledereli, Yilei Tu, Mrinmaya SachanICML 2025
