Text is Text, No Matter What: Unifying Text Recognition using Knowledge Distillation
Ayan Kumar Bhunia, Aneeshan Sain, Pinaki Nath Chowdhury, Yi-Zhe Song
摘要
Text recognition remains a fundamental and extensively researched topic in computer vision, largely owing to its wide array of commercial applications. The challenging nature of the very problem however dictated a fragmentation of research efforts: Scene Text Recognition (STR) that deals with text in everyday scenes, and Handwriting Text Recognition (HTR) that tackles hand-written text. In this paper, for the first time, we argue for their unification – we aim for a single model that can compete favourably with two separate state-of-the-art STR and HTR models. We first show that cross-utilisation of STR and HTR models trigger significant performance drops due to differences in their inherent challenges. We then tackle their union by introducing a knowledge distillation (KD) based framework. This however is non-trivial, largely due to the variable-length and sequential nature of text sequences, which renders off-the-shelf KD techniques that mostly work with global fixed length data, inadequate. For that, we propose four distillation losses, all of which are specifically designed to cope with the aforementioned unique characteristics of text recognition. Empirical evidence suggests that our proposed unified model performs at par with individual models, even surpassing them in certain cases. Ablative studies demonstrate that naive baselines such as a two-stage framework, multi-task and domain adaption/generalisation alternatives do not work that well, further authenticating our design.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper8
- Joint Visual Semantic Reasoning: Multi-Stage Decoder for Text RecognitionAyan Kumar Bhunia, Aneeshan Sain, Amandeep Kumar, Shuvozit Ghose 等ICCV 2021 · 被引用 60 次
- Self-supervised Character-to-Character Distillation for Text RecognitionTongkun Guan, Wei Shen, Xue Yang, Qi Feng 等ICCV 2023 · 被引用 36 次
- Self-Distillation Regularized Connectionist Temporal Classification Loss for Text Recognition: A Simple Yet Effective ApproachZiyin Zhang, Ning Lu, Minghui Liao, Yongshuai Huang 等AAAI 2024 · 被引用 20 次
- One-stage Low-resolution Text Recognition with High-resolution Knowledge TransferHang Guo, Tao Dai, Mingyan Zhu, Guanghao Meng 等ACM MM 2023 · 被引用 5 次
- Handwritten Text Generation from Visual ArchetypesVittorio Pippi, Silvia Cascianelli, Rita CucchiaraCVPR 2023
它引用的顶会 Paper15
- Improved Knowledge Distillation via Teacher AssistantSeyed-Iman Mirzadeh, Mehrdad Farajtabar, Ang Li, Nir Levine 等AAAI 2020 · 被引用 1,361 次
- Learning Lightweight Lane Detection CNNs by Self Attention DistillationYuenan Hou, Zheng Ma, Chunxiao Liu, Chen Change LoyICCV 2019 · 被引用 666 次
- Scene Text Visual Question AnsweringAli Furkan Biten, Rubèn Tito, Andrés Mafla, Lluís Gómez i Bigorda 等ICCV 2019 · 被引用 482 次
- Decoupled Attention Network for Text RecognitionTianwei Wang, Yuanzhi Zhu, Lianwen Jin, Canjie Luo 等AAAI 2020 · 被引用 289 次
- Relation Distillation Networks for Video Object DetectionJiajun Deng, Yingwei Pan, Ting Yao, Wengang Zhou 等ICCV 2019 · 被引用 211 次
相关 Paper
- Accurate Scene Text Recognition with Efficient Model Scaling and Cloze Self-DistillationAndrea Maracani, Savas Özkan, Sijun Cho, Hyowon Kim 等CVPR 2025
- Pushing the Performance Limit of Scene Text Recognizer without Human AnnotationCaiyuan Zheng, Hui Li, Seon-Min Rhee, Seungju Han 等CVPR 2022 · 被引用 20 次
- STIRER: A Unified Model for Low-Resolution Scene Text Image Recovery and RecognitionMinyi Zhao, Shijie Xuyang, Jihong Guan, Shuigeng ZhouACM MM 2023 · 被引用 9 次
- SwinTextSpotter: Scene Text Spotting via Better Synergy between Text Detection and Text RecognitionMingxin Huang, Yuliang Liu, Zhenghao Peng, Chongyu Liu 等CVPR 2022 · 被引用 150 次
- Out of Length Text Recognition with Sub-String MatchingYongkun Du, Zhineng Chen, Caiyan Jia, Xieping Gao 等AAAI 2025 · 被引用 9 次
