Text-Guided Explorable Image Super-Resolution
Kanchana Vaishnavi Gandikota, Paramanand Chandramouli
Abstract
In this paper, we introduce the problem of zero-shot textguided exploration of the solutions to open-domain image super-resolution. Our goal is to allow users to explore diverse, semantically accurate reconstructions that preserve data consistency with the low-resolution inputs for different large downsampling factors without explicitly training for these specific degradations. We propose two approaches for zero-shot text-guided super-resolution -i) modifying the generative process of text-to-image (T2I ) diffusion models to promote consistency with low-resolution inputs, and ii) incorporating language guidance into zero-shot diffusion-based restoration methods. We show that the proposed approaches result in diverse solutions that match the semantic meaning provided by the text prompt while preserving data consistency with the degraded inputs. We evaluate the proposed baselines for the task of extreme super-resolution and demonstrate advantages in terms of restoration quality, diversity, and explorability of solutions.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 6acb3eb2-7f5f-4331-b7ad-ce697d17afbcCited by top-tier papers4
- Self-Supervised Selective-Guided Diffusion Model for Old-Photo Face RestorationWenjie Li, Xiangyi Wang, Heng Guo, Guangwei Gao et al.NeurIPS 2025 · 13 citations
- LaSe-E2V: Towards Language-guided Semantic-aware Event-to-Video ReconstructionKanghao Chen, Hangyu Li, Jiazhou Zhou, Zeyu Wang et al.NeurIPS 2024 · 9 citations
- The Power of Context: How Multimodality Improves Image Super-ResolutionKangfu Mei, Hossein Talebi, Mojtaba Ardakani, Vishal M. Patel et al.CVPR 2025
- OVID: Open-Vocabulary Intrusion DetectionFujun Han, Jingqi Ye, Chenglong Zhang, Peng YeICLR 2026
Builds on43
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 13,211 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
Related papers
- Generative Powers of TenXiaojuan Wang, Janne Kontkanen, Brian Curless, Steven M. Seitz et al.CVPR 2024 · 3 citations
- Steered Diffusion: A Generalized Framework for Plug-and-Play Conditional Image SynthesisNithin Gopalakrishnan Nair, Anoop Cherian, Suhas Lohit, Ye Wang et al.ICCV 2023 · 22 citations
- TOSS: High-quality Text-guided Novel View Synthesis from a Single ImageYukai Shi, Jianan Wang, He Cao, Boshi Tang et al.ICLR 2024 · 28 citations
- Self-supervised Dynamic Heterogeneous Degradation Modeling for Unified Zero-Shot Image RestorationXiaowan Hu, Jing Yang, HeNan Liu, HuaQiu Li et al.CVPR 2026
- Harmonizing Visual and Textual Embeddings for Zero-Shot Text-to-Image CustomizationYeji Song, Jimyeong Kim, Wonhark Park, Wonsik Shin et al.AAAI 2025 · 6 citations
