Lune

ICML2026Top-tier venue

Seeing Symbols, Missing Structure: A Real-World Handwritten Mathematical Expression Recognition Benchmark for Large Models

Sheng Jiang, Lin Zhu, Runrui Li, Mei Wang, Qiannan Zhu, Yaoyao Zhong, Hua Huang

2026Year

Abstract

Handwritten mathematical expression recognition (HMER) remains challenging in real-world educational scenarios, even with recent advances in large vision-language models. While these models often achieve high accuracy in local symbol transcription, their reliability in capturing twodimensional mathematical structure under realistic handwritten conditions is still poorly understood. We introduce a real-world handwritten benchmark covering 13 categories of structurally complex expressions with authentic writing artifacts. Evaluations on large models reveal a clear performance degradation as structural complexity increases, even when symbol-level accuracy is high. Most failures arise from structural mis-parsing and context-dependent symbol role confusion rather than pure visual perception errors. To mitigate this issue, we propose a training-free, schema-anchored structure-aware inference framework that decomposes recognition into schema identification, schema-constrained transcription, and context-driven disambiguation.

Our method improves the ExpRate from 11.63% to 24.52% on Qwen-8B and generalizes well across multiple large models. Our benchmark provides a realistic evaluation for large models on handwritten mathematics, and our framework offers an effective and interpretable solution to structure-related failures in real-world HMER. Code and data are available at: https://github. com/BNU-ERC-ITEA/HMER-Bench.git.

Recent advances in encoder-decoder architectures and large vision-language models (VLMs) have substantially improved local symbol transcription accuracy (Yang et al., 2025b). However, in real educational scenarios, we observe a persistent gap between symbol-level correctness and expression-level validity. As illustrated in Fig. 1, models frequently produce outputs that contain mostly correct symbols but violate global structural constraints, such as mistaking single-line expressions as multi-line ones, confusing short division with long division, miscounting matrix size, or producing duplicated symbols, among others. These failures indicate that the dominant bottleneck of modern HMER systems lies not in visual perception, but in structural reasoning and global layout modeling.

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

lune papers fulltext d17a015d-a19f-4792-a045-427f39e10ae0

Builds on10

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines