Seeing Symbols, Missing Structure: A Real-World Handwritten Mathematical Expression Recognition Benchmark for Large Models
Sheng Jiang, Lin Zhu, Runrui Li, Mei Wang, Qiannan Zhu, Yaoyao Zhong, Hua Huang
Abstract
Handwritten mathematical expression recognition (HMER) remains challenging in real-world educational scenarios, even with recent advances in large vision-language models. While these models often achieve high accuracy in local symbol transcription, their reliability in capturing twodimensional mathematical structure under realistic handwritten conditions is still poorly understood. We introduce a real-world handwritten benchmark covering 13 categories of structurally complex expressions with authentic writing artifacts. Evaluations on large models reveal a clear performance degradation as structural complexity increases, even when symbol-level accuracy is high. Most failures arise from structural mis-parsing and context-dependent symbol role confusion rather than pure visual perception errors. To mitigate this issue, we propose a training-free, schema-anchored structure-aware inference framework that decomposes recognition into schema identification, schema-constrained transcription, and context-driven disambiguation.
Our method improves the ExpRate from 11.63% to 24.52% on Qwen-8B and generalizes well across multiple large models. Our benchmark provides a realistic evaluation for large models on handwritten mathematics, and our framework offers an effective and interpretable solution to structure-related failures in real-world HMER. Code and data are available at: https://github. com/BNU-ERC-ITEA/HMER-Bench.git.
Recent advances in encoder-decoder architectures and large vision-language models (VLMs) have substantially improved local symbol transcription accuracy (Yang et al., 2025b). However, in real educational scenarios, we observe a persistent gap between symbol-level correctness and expression-level validity. As illustrated in Fig. 1, models frequently produce outputs that contain mostly correct symbols but violate global structural constraints, such as mistaking single-line expressions as multi-line ones, confusing short division with long division, miscounting matrix size, or producing duplicated symbols, among others. These failures indicate that the dominant bottleneck of modern HMER systems lies not in visual perception, but in structural reasoning and global layout modeling.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d17a015d-a19f-4792-a045-427f39e10ae0Builds on10
- Is A Picture Worth A Thousand Words? Delving Into Spatial Reasoning for Vision Language ModelsJiayu Wang, Yifei Ming, Zhenmei Shi, Vibhav Vineet et al.NeurIPS 2024 · 166 citations
- Syntax-Aware Network for Handwritten Mathematical Expression RecognitionYe Yuan, Xiao Liu, Wondimu Dikubab, Hui Liu et al.CVPR 2022 · 74 citations
- What's "up" with vision-language models? Investigating their struggle with spatial reasoningAmita Kamath, Jack Hessel, Kai-Wei ChangEMNLP 2023 · 31 citations
- CC-OCR: A Comprehensive and Challenging OCR Benchmark for Evaluating Large Multimodal Models in LiteracyZhibo Yang, Jun Tang, Zhaohai Li, Pengfei Wang et al.ICCV 2025 · 15 citations
- Can Vision-Language Models Evaluate Handwritten Math?Oikantik Nath, Hanani Bathina, Mohammed Safi Ur Rahman Khan, Mitesh M. KhapraACL 2025 · 10 citations
Related papers
- Complex Mathematical Expression Recognition: Benchmark, Large-Scale Dataset and Strong BaselineWeikang Bai, Yongkun Du, Yuchen Su, Yazhen Xie et al.AAAI 2026 · 2 citations
- Uni-MuMER: Unified Multi-Task Fine-Tuning of Vision-Language Model for Handwritten Mathematical Expression RecognitionYu Li, Jin Jiang, Jianhua Zhu, Shuai Peng et al.NeurIPS 2025 · 7 citations
- TAMER: Tree-Aware Transformer for Handwritten Mathematical Expression RecognitionJianhua Zhu, Wenqi Zhao, Yu Li, Xingjian Hu et al.AAAI 2025 · 14 citations
- Read Ten Lines at One Glance: Line-Aware Semi-Autoregressive Transformer for Multi-Line Handwritten Mathematical Expression RecognitionWentao Yang, Zhe Li, Dezhi Peng, Lianwen Jin et al.ACM MM 2023 · 6 citations
- Art4Math: Handwritten Mathematical Expression Recognition via Multimodal Sketch GroundingYang Zhou, Jin Wang, Yuxiao Zhang, Kaixiang Huang et al.ACM MM 2025
