Lune

ICML2026顶会

Seeing Symbols, Missing Structure: A Real-World Handwritten Mathematical Expression Recognition Benchmark for Large Models

Sheng Jiang, Lin Zhu, Runrui Li, Mei Wang, Qiannan Zhu, Yaoyao Zhong, Hua Huang

出版方
2026年份

摘要

Handwritten mathematical expression recognition (HMER) remains challenging in real-world educational scenarios, even with recent advances in large vision-language models. While these models often achieve high accuracy in local symbol transcription, their reliability in capturing twodimensional mathematical structure under realistic handwritten conditions is still poorly understood. We introduce a real-world handwritten benchmark covering 13 categories of structurally complex expressions with authentic writing artifacts. Evaluations on large models reveal a clear performance degradation as structural complexity increases, even when symbol-level accuracy is high. Most failures arise from structural mis-parsing and context-dependent symbol role confusion rather than pure visual perception errors. To mitigate this issue, we propose a training-free, schema-anchored structure-aware inference framework that decomposes recognition into schema identification, schema-constrained transcription, and context-driven disambiguation.

Our method improves the ExpRate from 11.63% to 24.52% on Qwen-8B and generalizes well across multiple large models. Our benchmark provides a realistic evaluation for large models on handwritten mathematics, and our framework offers an effective and interpretable solution to structure-related failures in real-world HMER. Code and data are available at: https://github. com/BNU-ERC-ITEA/HMER-Bench.git.

Recent advances in encoder-decoder architectures and large vision-language models (VLMs) have substantially improved local symbol transcription accuracy (Yang et al., 2025b). However, in real educational scenarios, we observe a persistent gap between symbol-level correctness and expression-level validity. As illustrated in Fig. 1, models frequently produce outputs that contain mostly correct symbols but violate global structural constraints, such as mistaking single-line expressions as multi-line ones, confusing short division with long division, miscounting matrix size, or producing duplicated symbols, among others. These failures indicate that the dominant bottleneck of modern HMER systems lies not in visual perception, but in structural reasoning and global layout modeling.

问问这篇 Paper

智能体会读完全文。

Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。

可以从这些问题问起

智能体调用

Luneget_paper_fulltext

在 Lune 里问

免费开始,无需绑卡

lune papers fulltext d17a015d-a19f-4792-a045-427f39e10ae0

它引用的顶会 Paper10

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖