From Pixel to Precision: Enhancing Handwritten Mathematical Expression Recognition with Image-Level Reward
Ze Liu, Kai Zhang, Xianquan Wang, Shuochen Liu, Jiaxian Yan, Yupeng Han, Qi Liu
Abstract
Handwritten mathematical expression recognition is hindered by a fundamental misalignment between the dual representations of L A T E X formulas: the symbolic text and the rendered visual image. This discrepancy means that textually distinct L A T E X sequences can produce visually identical outputs, while minor textual errors can cause catastrophic rendering failures. As a result, text-level reward mechanisms cannot perfectly assess the quality of model predictions, failing to effectively guide the model towards optimal performance during training. To overcome this limitation, we introduce the Image Matching Score (IMS), a lightweight yet effective reward based on the structural edit distance of column-wise image projections, which robustly quantifies the visual fidelity between rendered formulas. Leveraging IMS, we then propose Image-Matching driven Policy Optimization (IMPO), a training framework built upon Group Relative Policy Optimization (GRPO). This approach facilitates stable policy learning directly from our sequence-level visual reward, notably without the need for a separate value function network. Extensive experiments demonstrate that IMPO yields consistent performance gains across various backbone models on the challenging CROHME, HME100K, and M 2 E datasets. Our framework establishes new state-of-the-art results, improving the Expression Recognition Rate by an average of 1.1% and up to 1.37% over strong prior methods.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 61b91905-c8ad-41cd-b302-ef3fad11fbaeBuilds on8
- Conservative Q-Learning for Offline Reinforcement LearningAviral Kumar, Aurick Zhou, George Tucker, Sergey LevineNeurIPS 2020 · 2,881 citations
- Syntax-Aware Network for Handwritten Mathematical Expression RecognitionYe Yuan, Xiao Liu, Wondimu Dikubab, Hui Liu et al.CVPR 2022 · 74 citations
- Handwritten Mathematical Expression Recognition via Attention Aggregation Based Bi-directional Mutual LearningXiaohang Bian, Bo Qin, Xiaozhe Xin, Jianwu Li et al.AAAI 2022 · 70 citations
- A Tree-Structured Decoder for Image-to-Markup GenerationJianshu Zhang, Jun Du, Yongxin Yang, Yi-Zhe Song et al.ICML 2020 · 68 citations
- TDv2: A Novel Tree-Structured Decoder for Offline Mathematical Expression RecognitionChangjie Wu, Jun Du, Yunqing Li, Jianshu Zhang et al.AAAI 2022 · 23 citations
Related papers
- Graph-to-Graph: Towards Accurate and Interpretable Online Handwritten Mathematical Expression RecognitionJin-Wen Wu, Fei Yin, Yan-Ming Zhang, Xu-Yao Zhang et al.AAAI 2021 · 39 citations
- Image Over Text: Transforming Formula Recognition Evaluation with Character Detection MatchingBin Wang, Fan Wu, Linke Ouyang, Zhuangcheng Gu et al.CVPR 2025
- Table2LaTeX-RL: High-Fidelity LaTeX Code Generation from Table Images via Reinforced Multimodal Language ModelsJun Ling, Yao Qi, Tao Huang, Shibo Zhou et al.NeurIPS 2025 · 9 citations
- Language Model is Suitable for Correction of Handwritten Mathematical Expressions RecognitionZui Chen, Jiaqi Han, Chaofan Yang, Yi ZhouEMNLP 2023 · 9 citations
- Structure-aware Mathematical Expression Recognition with Sequence-Level ModelingMinli Li, Peilin Zhao, Yifan Zhang, Shuaicheng Niu et al.ACM MM 2021 · 4 citations
