MathFlow: Enhancing the Perceptual Flow of MLLMs for Visual Mathematical Problems
Shuhang Chen, Hangjie Yuan, Yunqiu Xu, Pengwei Liu, Tao Feng, Jun Cen, Zeying Huang, Yi Yang
摘要
Despite strong results on many tasks, multimodal large language models (MLLMs) still underperform on visual mathematical problem solving, especially in reliably perceiving and interpreting diagrams. Inspired by human problem-solving, we hypothesize that the ability to extract meaningful information from diagrams is pivotal, as it directly conditions subsequent inference. Hence, we introduce FlowVerse, a comprehensive benchmark that provides a fine-grained evaluation of MLLMs' perception and reasoning capabilities. Our preliminary results on FlowVerse reveal that existing MLLMs exhibit substantial limitations when extracting essential information and reasoned properties from diagrams and performing complex reasoning based on these visual inputs. In response, we introduce MathFlow, a modular problemsolving pipeline that decouples perception and inference into distinct stages, thereby optimizing each independently. Given the perceptual limitations observed in current MLLMs, we trained MathFlow-P-7B as a dedicated perception model. Experimental results indicate that MathFlow-P-7B yields substantial performance gains when integrated with various closed-source and open-source inference models. This demonstrates the effectiveness of the MathFlow pipeline and its compatibility with diverse inference frameworks. Project page: https://github.com/MathFlow-zju/MathFlow .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- MC-Bench: A Benchmark for Multi-Context Visual Grounding in the Era of MLLMsYunqiu Xu, Linchao Zhu, Yi YangICCV 2025 · 被引用 7 次
- CogFlow: Bridging Perception and Reasoning through Knowledge Internalization for Visual Mathematical Problem SolvingShuhang Chen, Yunqiu Xu, Junjie Xie, Aojun Lu 等ICLR 2026 · 被引用 4 次
- MathSight: A Benchmark Exploring Have Vision-Language Models Really Seen in University-Level Mathematical Reasoning?Yuandong Wang, Yao Cui, Yuxin Zhao, Zhen Yang 等ACL 2026
- A Survey of Multimodal Mathematical Reasoning: From Perception, Alignment to ReasoningTianyu Yang, Sihong Wu, Yilun Zhao, Zhenwen Liang 等ACL 2026
它引用的顶会 Paper13
- MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual ContextsPan Lu, Hritik Bansal, Tony Xia, Jiacheng Liu 等ICLR 2024 · 被引用 1,472 次
- Overcoming Catastrophic Forgetting in Incremental Object Detection via Elastic Response DistillationTao Feng, Mang Wang, Hangjie YuanCVPR 2022 · 被引用 101 次
- Make Continual Learning Stronger via C-FlatAng Bian, Wei Li, Hangjie Yuan, Chengrong Yu 等NeurIPS 2024 · 被引用 48 次
- UniGeo: Unifying Geometry Logical Reasoning via Reformulating Mathematical ExpressionJiaqi Chen, Tong Li, Jinghui Qin, Pan Lu 等EMNLP 2022 · 被引用 37 次
- We-Math 2.0: A Versatile MathBook System for Incentivizing Visual Mathematical ReasoningRunqi Qiao, Qiuna Tan, Peiqing Yang, Yanzi Wang 等ICLR 2026 · 被引用 33 次
相关 Paper
- CircuitSense: A Hierarchical MLLM Benchmark Bridging Visual Comprehension and Symbolic Reasoning in Engineering Design ProcessArman Akbari, Jian Gao, Yifei Zou, Mei Yang 等ICLR 2026 · 被引用 3 次
- Math Blind: Failures in Diagram Understanding Undermine Reasoning in MLLMsYanpeng Sun, Shan Zhang, Wei Tang, Aotian Chen 等ICLR 2026 · 被引用 13 次
- Primitive Vision: Improving Diagram Understanding in MLLMsShan Zhang, Aotian Chen, Yanpeng Sun, Jindong Gu 等ICML 2025
- GNS: Solving Plane Geometry Problems by Neural-Symbolic Reasoning with Multi-Modal LLMsMaizhen Ning, Zihao Zhou, Qiufeng Wang, Xiaowei Huang 等AAAI 2025 · 被引用 10 次
- FlowGen: Synthesizing Diverse Flowcharts to Enhance and Benchmark MLLM ReasoningKaiwen Shi, Sichen Liu, Ziyue Lin, Hangrui Guo 等ICLR 2026
