Rethinking Super-Resolution as Text-Guided Details Generation
Chenxi Ma, Bo Yan, Qing Lin, Weimin Tan, Siming Chen
Abstract
Deep neural networks have greatly promoted the performance of single image super-resolution (SISR). Conventional methods still resort to restoring the single high-resolution (HR) solution only based on the input of image modality. However, the image-level information is insufficient to predict adequate details and photo-realistic visual quality facing large upscaling factors (×8, ×16). In this paper, we propose a new perspective that regards the SISR as a semantic image detail enhancement problem to generate semantically reasonable HR image that are faithful to the ground truth. To enhance the semantic accuracy and the visual quality of the reconstructed image, we explore the multi-modal fusion learning in SISR by proposing a Text-Guided Super-Resolution (TGSR) framework, which can effectively utilize the information from the text and image modalities. Different from existing methods, the proposed TGSR could generate HR image details that match the text descriptions through a coarse-to-fine process. Extensive experiments and ablation studies demonstrate the effect of the TGSR, which exploits the text reference to recover realistic images.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext afbe4fdb-eb3e-42ad-bd15-d229b32580c9Cited by top-tier papers2
- From Posterior Sampling to Meaningful Diversity in Image RestorationNoa Cohen, Hila Manor, Yuval Bahat, Tomer MichaeliICLR 2024 · 13 citations
- Text-Guided Explorable Image Super-ResolutionKanchana Vaishnavi Gandikota, Paramanand ChandramouliCVPR 2024
Builds on10
- Tag2Pix: Line Art Colorization Using Text Tag With SECat and Changing LossHyunsu Kim, Ho Young Jhoo, Eunhyeok Park, Sungjoo YooICCV 2019 · 119 citations
- ManiGAN: Text-Guided Image ManipulationBowen Li, Xiaojuan Qi, Thomas Lukasiewicz, Philip H. S. TorrCVPR 2020
- Deep Face Super-Resolution With Iterative Collaboration Between Attentive Recovery and Landmark EstimationCheng Ma, Zhenyu Jiang, Yongming Rao, Jiwen Lu et al.CVPR 2020
- Closed-Loop Matters: Dual Regression Networks for Single Image Super-ResolutionYong Guo, Jian Chen, Jingdong Wang, Qi Chen et al.CVPR 2020
- Structure-Preserving Super Resolution With Gradient GuidanceCheng Ma, Yongming Rao, Yean Cheng, Ce Chen et al.CVPR 2020
Related papers
- DF-GAN: A Simple and Effective Baseline for Text-to-Image SynthesisMing Tao, Hao Tang, Fei Wu, Xiaoyuan Jing et al.CVPR 2022 · 296 citations
- Text-Aware Real-World Image Super-Resolution via Diffusion Model with Joint Segmentation DecodersQiming Hu, Linlong Fan, Yiyan Luo, Yuhang Yu et al.NeurIPS 2025 · 7 citations
- Learning Texture Transformer Network for Image Super-ResolutionFuzhi Yang, Huan Yang, Jianlong Fu, Hongtao Lu et al.CVPR 2020
- Restore Text First, Enhance Image Later: Two-Stage Scene Text Image Super-Resolution with Glyph Structure GuidanceMinxing Luo, Linlong Fan, Qiushi Wang, Ge Wu et al.CVPR 2026 · 2 citations
- Text-IF: Leveraging Semantic Text Guidance for Degradation-Aware and Interactive Image FusionXunpeng Yi, Han Xu, Hao Zhang, Linfeng Tang et al.CVPR 2024 · 121 citations
