Complex Mathematical Expression Recognition: Benchmark, Large-Scale Dataset and Strong Baseline
Weikang Bai, Yongkun Du, Yuchen Su, Yazhen Xie, Zhineng Chen
摘要
Mathematical Expression Recognition (MER) has made significant progress in recognizing simple expressions, but the robust recognition of complex mathematical expressions with many tokens and multiple lines remains a formidable challenge. In this paper, we first introduce CMER-Bench, a carefully constructed benchmark that categorizes expressions into three difficulty levels: easy, moderate, and complex. Leveraging CMER-Bench, we conduct a comprehensive evaluation of existing MER models and general-purpose multimodal large language models (MLLMs). The results reveal that while current methods perform well on easy and moderate expressions, their performance degrades significantly when handling complex mathematical expressions, mainly because existing public training datasets are primarily composed of simple samples. In response, we propose MER-17M and CMER-3M that are large-scale datasets emphasizing the recognition of complex mathematical expressions. The datasets provide rich and diverse samples to support the development of accurate and robust complex MER models. Furthermore, to address the challenges posed by the complicated spatial layout of complex expressions, we introduce a novel expression tokenizer, and a new representation called Structured Mathematical Language, which explicitly models the hierarchical and spatial structure of expressions beyond LaTeX format. Based on these, we propose a specialized model named CMERNet, built upon an encoder-decoder architecture and trained on CMER-3M. Experimental results show that CMERNet, with only 125 million parameters, significantly outperforms existing MER models and MLLMs on CMER-Bench.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper8
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- CvT: Introducing Convolutions to Vision TransformersHaiping Wu, Bin Xiao, Noel Codella, Mengchen Liu 等ICCV 2021 · 被引用 2,397 次
- Patch n' Pack: NaViT, a Vision Transformer for any Aspect Ratio and ResolutionMostafa Dehghani, Basil Mustafa, Josip Djolonga, Jonathan Heek 等NeurIPS 2023 · 被引用 303 次
- Syntax-Aware Network for Handwritten Mathematical Expression RecognitionYe Yuan, Xiao Liu, Wondimu Dikubab, Hui Liu 等CVPR 2022 · 被引用 74 次
- TDv2: A Novel Tree-Structured Decoder for Offline Mathematical Expression RecognitionChangjie Wu, Jun Du, Yunqing Li, Jianshu Zhang 等AAAI 2022 · 被引用 23 次
相关 Paper
- UniMERNet: A Universal Network for Real-World Mathematical Expression RecognitionZhuangcheng Gu, Guang Liang, Bin Wang, Zhiyuan Zhao 等CVPR 2026
- Seeing Symbols, Missing Structure: A Real-World Handwritten Mathematical Expression Recognition Benchmark for Large ModelsSheng Jiang, Lin Zhu, Runrui Li, Mei Wang 等ICML 2026
- Speech-to-LaTeX: New Models and Datasets for Converting Spoken Equations and SentencesDmitrii Korzh, Dmitrii Tarasov, Artyom Iudin, Elvir Karimov 等ICLR 2026 · 被引用 2 次
- Structure-aware Mathematical Expression Recognition with Sequence-Level ModelingMinli Li, Peilin Zhao, Yifan Zhang, Shuaicheng Niu 等ACM MM 2021 · 被引用 4 次
- AffectGPT: A New Dataset, Model, and Benchmark for Emotion Understanding with Multimodal Large Language ModelsZheng Lian, Haoyu Chen, Lan Chen, Haiyang Sun 等ICML 2025
