VisioMath: Benchmarking Figure-based Mathematical Reasoning in LMMs
Can Li, Ying Liu, Ting Zhang, Mei Wang, Hua Huang
摘要
Large multimodal models have achieved remarkable progress in integrating vision and language, enabling strong performance across perception, reasoning, and domain-specific tasks. However, their capacity to reason over multiple, visually similar inputs remains insufficiently explored. Such fine-grained comparative reasoning is central to real-world tasks, especially in mathematics and education, where learners must often distinguish between nearly identical diagrams to identify correct solutions. To address this gap, we present VisioMath, a curated benchmark of 1,800 high-quality K-12 mathematics problems in which all candidate answers are diagrams with subtle visual similarities. A comprehensive evaluation of state-of-the-art LMMs, covering both leading closed-source systems and widely adopted open-source models, reveals a consistent decline in accuracy as inter-image similarity increases. Analysis indicates that the dominant failure mode stems from image-text misalignment: rather than grounding reasoning in textual cues, models often resort to shallow positional heuristics, resulting in systematic errors. We further explore three alignment-oriented strategies, spanning training-free approaches and finetuning, and achieve substantial accuracy gains. We hope that VisioMath will serve as a rigorous benchmark and catalyst for developing LMMs toward deeper diagram understanding, precise comparative reasoning, and grounded multi-image-text integration. The code and dataset are available at https://github.com/Nefefilibata/VisioMath .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- VisAidMath: Benchmarking Visual-Aided Mathematical ReasoningJingkun Ma, Runzhe Zhan, Yang Li, Di Sun 等ACL 2026 · 被引用 6 次
- Hierarchical Process Reward Models are Symbolic Vision LearnersShan Zhang, Aotian Chen, Kai Zou, Jindong Gu 等CVPR 2026 · 被引用 1 次
它引用的顶会 Paper11
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 被引用 11,349 次
- CogVLM: Visual Expert for Pretrained Language ModelsWeihan Wang, Qingsong Lv, Wenmeng Yu, Wenyi Hong 等NeurIPS 2024 · 被引用 858 次
- Vision-R1: Incentivizing Reasoning Capability in Multimodal Large Language ModelsWenxuan Huang, Bohan Jia, Shaosheng Cao, Zheyu Ye 等ICLR 2026 · 被引用 670 次
- MMMU: A Massive Multi-Discipline Multimodal Understanding and Reasoning Benchmark for Expert AGIXiang Yue, Yuansheng Ni, Tianyu Zheng, Kai Zhang 等CVPR 2024 · 被引用 213 次
- Math-PUMA: Progressive Upward Multimodal Alignment to Enhance Mathematical ReasoningWenwen Zhuang, Xin Huang, Xiantao Zhang, Jin ZengAAAI 2025 · 被引用 66 次
相关 Paper
- VisionMath: Vision-Form Mathematical Problem-SolvingZongyang Ma, Yuxin Chen, Ziqi Zhang, Zhongang Oi 等ICCV 2025 · 被引用 2 次
- MV-MATH: Evaluating Multimodal Math Reasoning in Multi-Visual ContextsPeijie Wang, Zhong-Zhi Li, Fei Yin, Dekang Ran 等CVPR 2025
- A Survey of Multimodal Mathematical Reasoning: From Perception, Alignment to ReasoningTianyu Yang, Sihong Wu, Yilun Zhao, Zhenwen Liang 等ACL 2026
- VisuLogic: A Benchmark for Evaluating Visual Reasoning in Multi-modal Large Language ModelsWeiye Xu, Jiahao Wang, Weiyun Wang, Zhe Chen 等ICLR 2026 · 被引用 103 次
- Integrating Visual Interpretation and Linguistic Reasoning for Geometric Problem SolvingZixian Guo, Ming Liu, Qilong Wang, Zhilong Ji 等ICCV 2025 · 被引用 1 次
