UniMERNet: A Universal Network for Real-World Mathematical Expression Recognition
Zhuangcheng Gu, Guang Liang, Bin Wang, Zhiyuan Zhao, Qintong Zhang, Weijia Li, Chao Xu, Bo Zhang, Botian Shi, Jiang Wu, Wentao Zhang, Conghui He
摘要
The paper introduces the UniMER dataset, marking the first study on Mathematical Expression Recognition (MER) targeting complex real-world scenarios. The UniMER dataset includes a large-scale training set, UniMER-1M, which offers unprecedented scale and diversity with one million training instances to train high-quality, robust models. Additionally, UniMER features a meticulously designed, diverse test set, UniMER-Test, which covers a variety of formula distributions found in real-world scenarios, providing a more comprehensive and fair evaluation. To better utilize the UniMER dataset, the paper proposes a Universal Mathematical Expression Recognition Network (UniMERNet), tailored to the characteristics of formula recognition. UniMERNet consists of a carefully designed encoder that incorporates detail-aware and local context features, and an optimized decoder for accelerated performance. Extensive experiments conducted using the UniMER-1M dataset and UniMERNet demonstrate that training on the large-scale UniMER-1M dataset can produce a more generalizable formula recognition model, significantly outperforming all previous datasets. Furthermore, the introduction of UniMERNet enhances the model's performance in formula recognition, achieving higher accuracy and speeds. All data, models, and code are available at https: //github.com/opendatalab/UniMERNet.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper12
- OCR Hinders RAG: Evaluating the Cascading Impact of OCR on Retrieval-Augmented GenerationJunyuan Zhang, Qintong Zhang, Bin Wang, Linke Ouyang 等ICCV 2025 · 被引用 10 次
- Towards Real-World Document Parsing via Realistic Scene Synthesis and Document-Aware TrainingGengluo Li, Pengyuan Lyu, Chengquan Zhang, Huawen Shen 等CVPR 2026 · 被引用 9 次
- Uni-MuMER: Unified Multi-Task Fine-Tuning of Vision-Language Model for Handwritten Mathematical Expression RecognitionYu Li, Jin Jiang, Jianhua Zhu, Shuai Peng 等NeurIPS 2025 · 被引用 7 次
- Doc-Researcher: A Unified System for Multimodal Document Parsing and Deep ResearchKuicai Dong, Shurui Huang, Fangda Ye, Wei Han 等WWW 2026 · 被引用 4 次
- Boosting Document Parsing Efficiency and Performance with Coarse-to-Fine Visual ProcessingCheng Cui, Ting Sun, Suyin Liang, Tingquan Gao 等CVPR 2026 · 被引用 3 次
它引用的顶会 Paper11
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 被引用 11,349 次
- MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language ModelsDeyao Zhu, Jun Chen, Xiaoqian Shen, Xiang Li 等ICLR 2024 · 被引用 3,079 次
- ConViT: Improving Vision Transformers with Soft Convolutional Inductive BiasesStéphane d'Ascoli, Hugo Touvron, Matthew L. Leavitt, Ari S. Morcos 等ICML 2021 · 被引用 1,021 次
- CMT: Convolutional Neural Networks Meet Vision TransformersJianyuan Guo, Kai Han, Han Wu, Yehui Tang 等CVPR 2022 · 被引用 839 次
- Conditional Positional Encodings for Vision TransformersXiangxiang Chu, Zhi Tian, Bo Zhang, Xinlong Wang 等ICLR 2023 · 被引用 406 次
相关 Paper
- Complex Mathematical Expression Recognition: Benchmark, Large-Scale Dataset and Strong BaselineWeikang Bai, Yongkun Du, Yuchen Su, Yazhen Xie 等AAAI 2026 · 被引用 2 次
- Syntax-Aware Network for Handwritten Mathematical Expression RecognitionYe Yuan, Xiao Liu, Wondimu Dikubab, Hui Liu 等CVPR 2022 · 被引用 74 次
- Seeing Symbols, Missing Structure: A Real-World Handwritten Mathematical Expression Recognition Benchmark for Large ModelsSheng Jiang, Lin Zhu, Runrui Li, Mei Wang 等ICML 2026
- Read Ten Lines at One Glance: Line-Aware Semi-Autoregressive Transformer for Multi-Line Handwritten Mathematical Expression RecognitionWentao Yang, Zhe Li, Dezhi Peng, Lianwen Jin 等ACM MM 2023 · 被引用 6 次
- TAMER: Tree-Aware Transformer for Handwritten Mathematical Expression RecognitionJianhua Zhu, Wenqi Zhao, Yu Li, Xingjian Hu 等AAAI 2025 · 被引用 14 次
