Event-Guided Scene Text Image Super-Resolution
Zihan Qi, Zeyu Xiao, Haoyi Zhao, Yang Zhao, Feng Xue, Wei Jia
Abstract
Scene text image super-resolution aims to enhance text legibility by recovering high-resolution text images from lowresolution inputs. However, maintaining fine details such as text strokes, edges, and textual accuracy remains challenging, particularly in low-light environments and high-speed motion scenarios, where degradation is more severe. Event cameras, with their high temporal resolution and ability to capture intensity changes, offer a promising solution for restoring lost fine details and mitigating degradation in these challenging conditions. In this paper, we propose EvTSR, the first framework that integrates Event data for scene Text image Super-Resolution. The core of EvTSR is the dualstream frequency boost (DSFB) mechanism, which separates image features into high-and low-frequency components. High-frequency details like edges and strokes are enhanced using event data via the event-guided high-frequency (EGH) mechanism, while low-frequency components, responsible for global structure, are refined using the Text-Guided Low-frequency (TGL) mechanism with a pre-trained text recognizer, ensuring textual coherence. To further improve cross-modal integration, we introduce the cross-modal fusion (CMF) mechanism, which effectively aligns event and image features, enabling robust information fusion. Extensive experiments demonstrate that EvTSR achieves superior performance over existing methods.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext a793a639-e002-4e79-88fc-fb68ee644173Builds on22
- Toward Real-World Single Image Super-Resolution: A New Benchmark and a New ModelJianrui Cai, Hui Zeng, Hongwei Yong, Zisheng Cao et al.ICCV 2019 · 713 citations
- A Text Attention Network for Spatial Deformation Robust Scene Text Image Super-resolutionJianqi Ma, Zhetong Liang, Lei ZhangCVPR 2022 · 95 citations
- Spatio-Temporal Recurrent Networks for Event-Based Optical Flow EstimationZiluo Ding, Rui Zhao, Jiyuan Zhang, Tianxiao Gao et al.AAAI 2022 · 76 citations
- Text Gestalt: Stroke-Aware Scene Text Image Super-resolutionJingye Chen, Haiyang Yu, Jianqi Ma, Bin Li et al.AAAI 2022 · 63 citations
- Scene Text Image Super-Resolution via Parallelly Contextual Attention NetworkCairong Zhao, Shuyang Feng, Brian Nlong Zhao, Zhijun Ding et al.ACM MM 2021 · 61 citations
Related papers
- EvIntSR-Net: Event Guided Multiple Latent Frames Reconstruction and Super-resolutionJin Han, Yixin Yang, Chu Zhou, Chao Xu et al.ICCV 2021 · 57 citations
- Exploiting Blurry Representations for Event-guided Video Super-ResolutionZeyu Xiao, Xinchao WangAAAI 2026
- EvTexture: Event-driven Texture Enhancement for Video Super-ResolutionDachun Kai, Jiayao Lu, Yueyi Zhang, Xiaoyan SunICML 2024 · 14 citations
- Frequency-Aware Event-Based Video Deblurring for Real-World Motion BlurTaewoo Kim, Hoonhee Cho, Kuk-Jin YoonCVPR 2024
- Event-based Motion Deblurring with Modality-Aware Decomposition and RecompositionWen Yang, Jinjian Wu, Leida Li, Weisheng Dong et al.ACM MM 2023 · 14 citations
