Lune

AAAI2026Top-tier venue

Event-Guided Scene Text Image Super-Resolution

Zihan Qi, Zeyu Xiao, Haoyi Zhao, Yang Zhao, Feng Xue, Wei Jia

2026Year

Abstract

Scene text image super-resolution aims to enhance text legibility by recovering high-resolution text images from lowresolution inputs. However, maintaining fine details such as text strokes, edges, and textual accuracy remains challenging, particularly in low-light environments and high-speed motion scenarios, where degradation is more severe. Event cameras, with their high temporal resolution and ability to capture intensity changes, offer a promising solution for restoring lost fine details and mitigating degradation in these challenging conditions. In this paper, we propose EvTSR, the first framework that integrates Event data for scene Text image Super-Resolution. The core of EvTSR is the dualstream frequency boost (DSFB) mechanism, which separates image features into high-and low-frequency components. High-frequency details like edges and strokes are enhanced using event data via the event-guided high-frequency (EGH) mechanism, while low-frequency components, responsible for global structure, are refined using the Text-Guided Low-frequency (TGL) mechanism with a pre-trained text recognizer, ensuring textual coherence. To further improve cross-modal integration, we introduce the cross-modal fusion (CMF) mechanism, which effectively aligns event and image features, enabling robust information fusion. Extensive experiments demonstrate that EvTSR achieves superior performance over existing methods.

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

lune papers fulltext a793a639-e002-4e79-88fc-fb68ee644173

Builds on22

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines