Math Blind: Failures in Diagram Understanding Undermine Reasoning in MLLMs
Yanpeng Sun, Shan Zhang, Wei Tang, Aotian Chen, Piotr Koniusz, Kai Zou, Yuan Xue, Anton van den Hengel
摘要
Diagrams represent a form of visual language that encodes abstract concepts and relationships through structured symbols and their spatial arrangements. Unlike natural images, they are inherently symbolic, and entirely artificial. They thus pose unique challenges for Multimodal Large Language Models (MLLMs) distinct from natural image processing. Recent studies have shown that MLLMs often exhibit flawed reasoning and hallucinations when handling diagram inputs. We investigate here whether these limitations stem from shortcomings in the models' ability to interpret diagrams themselves. To this end, we develop a diagnostic test suite that isolates perception from reasoning. Our systematic evaluation reveals that MLLMs perform poorly on basic perceptual tasks, e.g., shape classification, object counting, relationship identification, and object grounding, with near-zero accuracy on fine-grained grounding. Further analysis shows that weak diagram perception leads to ``blind faith in text", where models rely on textual shortcuts rather than visual understanding (that is, they are ). We hypothesize that enabling models to capture the inherent structural properties of diagrams, represented as graphs of primitives and their interrelationships, is essential for improving diagram understanding. Experiments with 7B and 32B MLLMs validate this assumption, with models trained on such representations achieving a +79% gain on the grounding task. Crucially, these gains transfer to reasoning, achieving 3–4% cross-suite improvements on four public benchmarks even without additional chain-of-thought reasoning data. Our findings demonstrate that low-level perception supports faithful high-level reasoning in mathematical MLLMs. We provide both methodological frameworks and empirical evidence to guide future research in this direction. Our project page is at blueviocean/.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- A Survey of Deep Learning for Geometry Problem SolvingJianzhe Ma, Wenxuan Wang, Qin JinACL 2026 · 被引用 5 次
- GeoTikzBridge: Advancing Multimodal Code Generation for Geometric Perception and ReasoningJiayin Sun, Caixia Sun, Boyu Yang, hailin li 等CVPR 2026 · 被引用 3 次
- Hierarchical Process Reward Models are Symbolic Vision LearnersShan Zhang, Aotian Chen, Kai Zou, Jindong Gu 等CVPR 2026 · 被引用 1 次
- Hierarchically Robust Zero-shot Vision-language ModelsJunhao Dong, Yifei Zhang, Hao Zhu, Yew-Soon Ong 等CVPR 2026 · 被引用 1 次
它引用的顶会 Paper25
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 被引用 11,349 次
- Deformable DETR: Deformable Transformers for End-to-End Object DetectionXizhou Zhu, Weijie Su, Lewei Lu, Bin Li 等ICLR 2021 · 被引用 7,353 次
- Flamingo: a Visual Language Model for Few-Shot LearningJean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech 等NeurIPS 2022 · 被引用 6,707 次
- PaLM-E: An Embodied Multimodal Language ModelDanny Driess, Fei Xia, Mehdi S. M. Sajjadi, Corey Lynch 等ICML 2023 · 被引用 2,601 次
- MM-Vet: Evaluating Large Multimodal Models for Integrated CapabilitiesWeihao Yu, Zhengyuan Yang, Linjie Li, Jianfeng Wang 等ICML 2024 · 被引用 1,191 次
相关 Paper
- Do Vision-Language Models Really Understand Visual Language?Yifan Hou, Buse Giledereli, Yilei Tu, Mrinmaya SachanICML 2025
- Primitive Vision: Improving Diagram Understanding in MLLMsShan Zhang, Aotian Chen, Yanpeng Sun, Jindong Gu 等ICML 2025
- MathFlow: Enhancing the Perceptual Flow of MLLMs for Visual Mathematical ProblemsShuhang Chen, Hangjie Yuan, Yunqiu Xu, Pengwei Liu 等ACL 2026 · 被引用 9 次
- GlFoMR: A Glance-then-Focus Multimodal Reasoning Framework for Diagram Question AnsweringYaxian Wang, Bifan Wei, Jun Liu, Lingling Zhang 等SIGIR 2025
- CircuitSense: A Hierarchical MLLM Benchmark Bridging Visual Comprehension and Symbolic Reasoning in Engineering Design ProcessArman Akbari, Jian Gao, Yifei Zou, Mei Yang 等ICLR 2026 · 被引用 3 次
