The Power of Context: How Multimodality Improves Image Super-Resolution
Kangfu Mei, Hossein Talebi, Mojtaba Ardakani, Vishal M. Patel, Peyman Milanfar, Mauricio Delbracio
Abstract
Inputs Outputs Reference A close-up of a male lion with a dark mane, light tan face, and pink tongue sticking out . . . LR LR (Zoomed) Caption PASD SeeSR MMSR (Ours) HR Depth Segmentation Edge PASD (Zoomed) SeeSR (Zoomed) MMSR (Zoomed) HR (Zoomed) Figure 1. Our Multimodal Super-Resolution (MMSR) method leverages the rich context of multimodal guidance, including image captions, depth maps, semantic segmentation maps, and edges inferred from LR. MMSR surpasses state-of-the-art methods by producing more realistic results and suppressing artifacts that, while plausible, are inconsistent with the information present in the LR input.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 0d180834-0a9a-4259-8d73-4b5f4ffe4936Cited by top-tier papers8
- 4KAgent: Agentic Any Image to 4K Super-ResolutionYushen Zuo, Qi Zheng, Mingyang Wu, Xinrui Jiang et al.NeurIPS 2025 · 51 citations
- Exploiting Diffusion Prior for Task-Driven Image RestorationJaeha Kim, Junghun Oh, Kyoung Mu LeeICCV 2025 · 6 citations
- Kernel Density Steering: Inference-Time Scaling via Mode Seeking for Image RestorationYuyang Hu, Kangfu Mei, Mojtaba Sahraee-Ardakan, Ulugbek Kamilov et al.NeurIPS 2025 · 6 citations
- From Inpainting to Layer Decomposition: Repurposing Generative Inpainting Models for Image Layer DecompositionJingxi Chen, Yixiao Zhang, Xiaoye Qian, Zongxia Li et al.CVPR 2026 · 5 citations
- Fine-Structure Preserved Real-World Image Super-Resolution Via Transfer Vae TrainingQiaosi Yi, Shuai Liu, Rongyuan Wu, Lingchen Sun et al.ICCV 2025 · 4 citations
Builds on40
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 13,211 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
- Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingChitwan Saharia, William Chan, Saurabh Saxena, Lala Li et al.NeurIPS 2022 · 8,965 citations
Related papers
- Rethinking Super-Resolution as Text-Guided Details GenerationChenxi Ma, Bo Yan, Qing Lin, Weimin Tan et al.ACM MM 2022 · 6 citations
- GLEAN: Generative Latent Bank for Large-Factor Image Super-ResolutionKelvin C. K. Chan, Xintao Wang, Xiangyu Xu, Jinwei Gu et al.CVPR 2021
- Perceive, Understand and Restore: Real-World Image Super-Resolution with Autoregressive Multimodal Generative ModelsHongyang Wei, Shuaizheng Liu, Chun Yuan, Lei ZhangICCV 2025 · 4 citations
- SeD: Semantic-Aware Discriminator for Image Super-ResolutionBingchen Li, Xin Li, Hanxin Zhu, Yeying Jin et al.CVPR 2024
- Restore Text First, Enhance Image Later: Two-Stage Scene Text Image Super-Resolution with Glyph Structure GuidanceMinxing Luo, Linlong Fan, Qiushi Wang, Ge Wu et al.CVPR 2026 · 2 citations
