Seeing Symbols, Missing Structure: A Real-World Handwritten Mathematical Expression Recognition Benchmark for Large Models
Sheng Jiang, Lin Zhu, Runrui Li, Mei Wang, Qiannan Zhu, Yaoyao Zhong, Hua Huang
摘要
Handwritten mathematical expression recognition (HMER) remains challenging in real-world educational scenarios, even with recent advances in large vision-language models. While these models often achieve high accuracy in local symbol transcription, their reliability in capturing twodimensional mathematical structure under realistic handwritten conditions is still poorly understood. We introduce a real-world handwritten benchmark covering 13 categories of structurally complex expressions with authentic writing artifacts. Evaluations on large models reveal a clear performance degradation as structural complexity increases, even when symbol-level accuracy is high. Most failures arise from structural mis-parsing and context-dependent symbol role confusion rather than pure visual perception errors. To mitigate this issue, we propose a training-free, schema-anchored structure-aware inference framework that decomposes recognition into schema identification, schema-constrained transcription, and context-driven disambiguation.
Our method improves the ExpRate from 11.63% to 24.52% on Qwen-8B and generalizes well across multiple large models. Our benchmark provides a realistic evaluation for large models on handwritten mathematics, and our framework offers an effective and interpretable solution to structure-related failures in real-world HMER. Code and data are available at: https://github. com/BNU-ERC-ITEA/HMER-Bench.git.
Recent advances in encoder-decoder architectures and large vision-language models (VLMs) have substantially improved local symbol transcription accuracy (Yang et al., 2025b). However, in real educational scenarios, we observe a persistent gap between symbol-level correctness and expression-level validity. As illustrated in Fig. 1, models frequently produce outputs that contain mostly correct symbols but violate global structural constraints, such as mistaking single-line expressions as multi-line ones, confusing short division with long division, miscounting matrix size, or producing duplicated symbols, among others. These failures indicate that the dominant bottleneck of modern HMER systems lies not in visual perception, but in structural reasoning and global layout modeling.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper10
- Is A Picture Worth A Thousand Words? Delving Into Spatial Reasoning for Vision Language ModelsJiayu Wang, Yifei Ming, Zhenmei Shi, Vibhav Vineet 等NeurIPS 2024 · 被引用 166 次
- Syntax-Aware Network for Handwritten Mathematical Expression RecognitionYe Yuan, Xiao Liu, Wondimu Dikubab, Hui Liu 等CVPR 2022 · 被引用 74 次
- What's "up" with vision-language models? Investigating their struggle with spatial reasoningAmita Kamath, Jack Hessel, Kai-Wei ChangEMNLP 2023 · 被引用 31 次
- CC-OCR: A Comprehensive and Challenging OCR Benchmark for Evaluating Large Multimodal Models in LiteracyZhibo Yang, Jun Tang, Zhaohai Li, Pengfei Wang 等ICCV 2025 · 被引用 15 次
- Can Vision-Language Models Evaluate Handwritten Math?Oikantik Nath, Hanani Bathina, Mohammed Safi Ur Rahman Khan, Mitesh M. KhapraACL 2025 · 被引用 10 次
相关 Paper
- Complex Mathematical Expression Recognition: Benchmark, Large-Scale Dataset and Strong BaselineWeikang Bai, Yongkun Du, Yuchen Su, Yazhen Xie 等AAAI 2026 · 被引用 2 次
- Uni-MuMER: Unified Multi-Task Fine-Tuning of Vision-Language Model for Handwritten Mathematical Expression RecognitionYu Li, Jin Jiang, Jianhua Zhu, Shuai Peng 等NeurIPS 2025 · 被引用 7 次
- TAMER: Tree-Aware Transformer for Handwritten Mathematical Expression RecognitionJianhua Zhu, Wenqi Zhao, Yu Li, Xingjian Hu 等AAAI 2025 · 被引用 14 次
- Read Ten Lines at One Glance: Line-Aware Semi-Autoregressive Transformer for Multi-Line Handwritten Mathematical Expression RecognitionWentao Yang, Zhe Li, Dezhi Peng, Lianwen Jin 等ACM MM 2023 · 被引用 6 次
- Art4Math: Handwritten Mathematical Expression Recognition via Multimodal Sketch GroundingYang Zhou, Jin Wang, Yuxiao Zhang, Kaixiang Huang 等ACM MM 2025
