Pixel Adapter: A Graph-Based Post-Processing Approach for Scene Text Image Super-Resolution
Wenyu Zhang, Xin Deng, Baojun Jia, Xingtong Yu, Yifan Chen, Jin Ma, Qing Ding, Xinming Zhang
Abstract
Current Scene text image super-resolution approaches primarily focus on extracting robust features, acquiring text information, and complex training strategies to generate super-resolution images. However, the upsampling module, which is crucial in the process of converting low-resolution images to high-resolution ones, has received little attention in existing works. To address this issue, we propose the Pixel Adapter Module (PAM) based on graph attention to address pixel distortion caused by upsampling. The PAM effectively captures local structural information by allowing each pixel to interact with its neighbors and update features. Unlike previous graph attention mechanisms, our approach achieves 2-3 orders of magnitude improvement in efficiency and memory utilization by eliminating the dependency on sparse adjacency matrices and introducing a sliding window approach for efficient parallel computation. Additionally, we introduce the MLP-based Sequential Residual Block (MSRB) for robust feature extraction from text images, and a Local Contour Awareness loss (L 𝑙𝑐𝑎 ) to enhance the model's perception of details. Comprehensive experiments on TextZoom demonstrate that our proposed method generates highquality super-resolution images, surpassing existing methods in recognition accuracy. For single-stage and multi-stage strategies, we achieved improvements of 0.7% and 2.6%, respectively, increasing the performance from 52.6% and 53.7% to 53.3% and 56.3%. The code is available at https://github.com/wenyu1009/RTSRN.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 9c4b2501-025b-48a6-83db-afed8831eee4Cited by top-tier papers5
- PEAN: A Diffusion-Based Prior-Enhanced Attention Network for Scene Text Image Super-ResolutionZuoyan Zhao, Hui Xue, Pengfei Fang, Shipeng ZhuACM MM 2024 · 16 citations
- GlyphSR: A Simple Glyph-Aware Framework for Scene Text Image Super-ResolutionBaole Wei, Yuxuan Zhou, Liangcai Gao, Zhi TangAAAI 2025 · 4 citations
- StyleSRN: Scene Text Image Super-Resolution with Text Style EmbeddingShengrong Yuan, Runmin Wang, Ke Hao, Xuqi Ma et al.ICCV 2025 · 2 citations
- Suppressing Uncertainties in Degradation Estimation for Blind Super-ResolutionJunxiong Lin, Zen Tao, Xuan Tong, Xinji Mai et al.ACM MM 2024 · 2 citations
- Number it: Temporal Grounding Videos like Flipping MangaYongliang Wu, Xinting Hu, Yuyang Sun, Yizhou Zhou et al.CVPR 2025
Builds on11
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- A ConvNet for the 2020sZhuang Liu, Hanzi Mao, Chao-Yuan Wu, Christoph Feichtenhofer et al.CVPR 2022 · 6,782 citations
- MLP-Mixer: An all-MLP Architecture for VisionIlya O. Tolstikhin, Neil Houlsby, Alexander Kolesnikov, Lucas Beyer et al.NeurIPS 2021 · 3,862 citations
- Toward Real-World Single Image Super-Resolution: A New Benchmark and a New ModelJianrui Cai, Hui Zeng, Hongwei Yong, Zisheng Cao et al.ICCV 2019 · 713 citations
Related papers
- Scene Text Image Super-Resolution via Parallelly Contextual Attention NetworkCairong Zhao, Shuyang Feng, Brian Nlong Zhao, Zhijun Ding et al.ACM MM 2021 · 61 citations
- Gradient-Based Graph Attention for Scene Text Image Super-resolutionXiangyuan Zhu, Kehua Guo, Hui Fang, Rui Ding et al.AAAI 2023 · 18 citations
- A Text Attention Network for Spatial Deformation Robust Scene Text Image Super-resolutionJianqi Ma, Zhetong Liang, Lei ZhangCVPR 2022 · 95 citations
- Scene Text Telescope: Text-Focused Scene Image Super-ResolutionJingye Chen, Bin Li, Xiangyang XueCVPR 2021
- Text Gestalt: Stroke-Aware Scene Text Image Super-resolutionJingye Chen, Haiyang Yu, Jianqi Ma, Bin Li et al.AAAI 2022 · 63 citations
