VisionMath: Vision-Form Mathematical Problem-Solving
Zongyang Ma, Yuxin Chen, Ziqi Zhang, Zhongang Oi, Chunfeng Yuan, Shaojie Zhu, Chengxiang Zhuo, Bing Li, Ye Liu, Zang Li, Ying Shan, Weiming Hu
Abstract
Mathematical problems in real-world scenarios are often presented in a purely vision-form, where textual problem statement and accompanying math figures, e.g., geometry figures and functional graphs, are integrated into a single image. This vision-form problem-solving task requires precise comprehension and reasoning on both textual and graphical elements in the images, posing significant challenge to current Multimodal Large Language Models (MLLMs), which process text and math figures in isolation. In this work, we propose Vision-Math, the first exploration for vision-form mathematical problem-solving model, which employs a three-stage progressive multimodal reasoning alignment strategy to systematically enhance task-specific capabilities. Building upon a LLM proficient in unimodal mathematical reasoning, VisionMath first establishes foundational OCR capabilities through capturing rendered mathematical problem images. Subsequently, the model develops comprehensive understanding of figure structures and properties via learning from figure descriptions and mathematical educational videos. Finally, the model's reasoning capacity is activated using carefully constructed visual-form problemsolving datasets VisionMath-IT with chain-of-thought annotations. For comprehensive evaluation, we construct multilingual benchmarks covering diverse problem types, including geometry, algebra, function problems in both English and Chinese. Experimental results demonstrate that Vision-Math significantly outperforms existing general-purpose
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext db304598-37ac-40ef-9359-ad4bee944073Cited by top-tier papers3
- UniPixel: Unified Object Referring and Segmentation for Pixel-Level Visual ReasoningYe Liu, Zongyang Ma, Junfu Pu, Zhongang Qi et al.NeurIPS 2025 · 39 citations
- VideoMind: A Chain-of-LoRA Agent for Temporal-Grounded Video ReasoningYe Liu, Kevin Qinghong Lin, Chang Wen Chen, Mike Zheng ShouICLR 2026 · 23 citations
- STAR-R1: Multi-View Spatial TrAnsformation Reasoning by Reinforcing Multimodal LLMsZongzhao Li, Zongyang Ma, Mingze Li, Songyou Li et al.CVPR 2026
Builds on14
- TURL: Table Understanding through Representation LearningXiang Deng, Huan Sun, Alyssa Lees, You Wu et al.VLDB 2021 · 2,406 citations
- MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual ContextsPan Lu, Hritik Bansal, Tony Xia, Jiacheng Liu et al.ICLR 2024 · 1,472 citations
- TabFact: A Large-scale Dataset for Table-based Fact VerificationWenhu Chen, Hongmin Wang, Jianshu Chen, Yunkai Zhang et al.ICLR 2020 · 674 citations
- MMMU: A Massive Multi-Discipline Multimodal Understanding and Reasoning Benchmark for Expert AGIXiang Yue, Yuansheng Ni, Tianyu Zheng, Kai Zhang et al.CVPR 2024 · 213 citations
- VisualMRC: Machine Reading Comprehension on Document ImagesRyota Tanaka, Kyosuke Nishida, Sen YoshidaAAAI 2021 · 201 citations
Related papers
- VisioMath: Benchmarking Figure-based Mathematical Reasoning in LMMsCan Li, Ying Liu, Ting Zhang, Mei Wang et al.ICLR 2026 · 7 citations
- Integrating Visual Interpretation and Linguistic Reasoning for Geometric Problem SolvingZixian Guo, Ming Liu, Qilong Wang, Zhilong Ji et al.ICCV 2025 · 1 citation
- MV-MATH: Evaluating Multimodal Math Reasoning in Multi-Visual ContextsPeijie Wang, Zhong-Zhi Li, Fei Yin, Dekang Ran et al.CVPR 2025
- MathSight: A Benchmark Exploring Have Vision-Language Models Really Seen in University-Level Mathematical Reasoning?Yuandong Wang, Yao Cui, Yuxin Zhao, Zhen Yang et al.ACL 2026
- A Survey of Multimodal Mathematical Reasoning: From Perception, Alignment to ReasoningTianyu Yang, Sihong Wu, Yilun Zhao, Zhenwen Liang et al.ACL 2026
