UPOCR: Towards Unified Pixel-Level OCR Interface
Dezhi Peng, Zhenhua Yang, Jiaxin Zhang, Chongyu Liu, Yongxin Shi, Kai Ding, Fengjun Guo, Lianwen Jin
摘要
In recent years, the optical character recognition (OCR) field has been proliferating with plentiful cutting-edge approaches for a wide spectrum of tasks. However, these approaches are task-specifically designed with divergent paradigms, architectures, and training strategies, which significantly increases the complexity of research and maintenance and hinders the fast deployment in applications. To this end, we propose UPOCR, a simple-yet-effective generalist model for Unified Pixel-level OCR interface. Specifically, the UPOCR unifies the paradigm of diverse OCR tasks as image-to-image transformation and the architecture as a vision Transformer (ViT)-based encoder-decoder. Learnable task prompts are introduced to push the general feature representations extracted by the encoder toward task-specific spaces, endowing the decoder with task awareness. Moreover, the model training is uniformly aimed at minimizing the discrepancy between the generated and ground-truth images regardless of the inhomogeneity among tasks. Experiments are conducted on three pixel-level OCR tasks including text removal, text segmentation, and tampered text detection. Without bells and whistles, the experimental results showcase that the proposed method can simultaneously achieve state-of-the-art performance on three tasks with a unified single model, which provides valuable strategies and insights for future research on generalist OCR models. Code will be publicly available.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Predicting the Original Appearance of Damaged Historical DocumentsZhenhua Yang, Dezhi Peng, Yongxin Shi, Yuyi Zhang 等AAAI 2025 · 被引用 8 次
- InstructOCR: Instruction Boosting Scene Text SpottingChen Duan, Qianyi Jiang, Pei Fu, Jiamin Chen 等AAAI 2025 · 被引用 7 次
- DocSAM: Unified Document Image Segmentation via Query Decomposition and Heterogeneous Mixed LearningXiao-Hui Li, Fei Yin, Cheng-Lin LiuCVPR 2025
- ST-SAM: Multimodal Scene Text Segmentation with Dense Visual and Sparse Textual Prompts via SAMJin Wei, Yaqiang Wu, Jiayi Yan, Zeng Li 等AAAI 2026
它引用的顶会 Paper31
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu 等ICCV 2021 · 被引用 31,683 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 被引用 11,349 次
- SegFormer: Simple and Efficient Design for Semantic Segmentation with TransformersEnze Xie, Wenhai Wang, Zhiding Yu, Anima Anandkumar 等NeurIPS 2021 · 被引用 9,661 次
相关 Paper
- ViTEraser: Harnessing the Power of Vision Transformers for Scene Text Removal with SegMIM PretrainingDezhi Peng, Chongyu Liu, Yuliang Liu, Lianwen JinAAAI 2024 · 被引用 18 次
- UNIT: Unifying Image and Text Recognition in One Vision EncoderYi Zhu, Yanpeng Zhou, Chunwei Wang, Yang Cao 等NeurIPS 2024 · 被引用 15 次
- TrOCR: Transformer-Based Optical Character Recognition with Pre-trained ModelsMinghao Li, Tengchao Lv, Jingye Chen, Lei Cui 等AAAI 2023 · 被引用 607 次
- ODM: A Text-Image Further Alignment Pre-training Approach for Scene Text Detection and SpottingChen Duan, Pei Fu, Shan Guo, Qianyi Jiang 等CVPR 2024
- Uni-Perceiver v2: A Generalist Model for Large-Scale Vision and Vision-Language TasksHao Li, Jinguo Zhu, Xiaohu Jiang, Xizhou Zhu 等CVPR 2023
