Iterative Token Evaluation and Refinement for Real-World Super-resolution
Chaofeng Chen, Shangchen Zhou, Liang Liao, Haoning Wu, Wenxiu Sun, Qiong Yan, Weisi Lin
Abstract
Real-world image super-resolution (RWSR) is a long-standing problem as low-quality (LQ) images often have complex and unidentified degradations. Existing methods such as Generative Adversarial Networks (GANs) or continuous diffusion models present their own issues including GANs being difficult to train while continuous diffusion models requiring numerous inference steps. In this paper, we propose an Iterative Token Evaluation and Refinement (ITER) framework for RWSR, which utilizes a discrete diffusion model operating in the discrete token representation space, i.e., indexes of features extracted from a VQGAN codebook pre-trained with high-quality (HQ) images. We show that ITER is easier to train than GANs and more efficient than continuous diffusion models. Specifically, we divide RWSR into two sub-tasks, i.e., distortion removal and texture generation. Distortion removal involves simple HQ token prediction with LQ images, while texture generation uses a discrete diffusion model to iteratively refine the distortion removal output with a token refinement network. In particular, we propose to include a token evaluation network in the discrete diffusion process. It learns to evaluate which tokens are good restorations and helps to improve the iterative refinement results. Moreover, the evaluation network can first check status of the distortion removal output and then adaptively select total refinement steps needed, thereby maintaining a good balance between distortion removal and texture generation. Extensive experimental results show that ITER is easy to train and performs well within just 8 iterative steps.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers4
- AI-generated Image Quality Assessment in Visual CommunicationYu Tian, Yixuan Li, Baoliang Chen, Hanwei Zhu et al.AAAI 2025 · 12 citations
- SeeClear: Semantic Distillation Enhances Pixel Condensation for Video Super-ResolutionQi Tang, Yao Zhao, Meiqin Liu, Chao YaoNeurIPS 2024 · 10 citations
- 3DEnhancer: Consistent Multi-View Diffusion for 3D EnhancementYihang Luo, Shangchen Zhou, Yushi Lan, Xingang Pan et al.CVPR 2025
- Iterative Predictor-Critic Code Decoding for Real-World Image DehazingJiayi Fu, Siyu Liu, Zikun Liu, Chun-Le Guo et al.CVPR 2025
Builds on20
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Restormer: Efficient Transformer for High-Resolution Image RestorationSyed Waqas Zamir, Aditya Arora, Salman Khan, Munawar Hayat et al.CVPR 2022 · 3,348 citations
- Vector Quantized Diffusion Model for Text-to-Image SynthesisShuyang Gu, Dong Chen, Jianmin Bao, Fang Wen et al.CVPR 2022 · 607 citations
- Towards Robust Blind Face Restoration with Codebook Lookup TransformerShangchen Zhou, Kelvin C. K. Chan, Chongyi Li, Chen Change LoyNeurIPS 2022 · 431 citations
Related papers
- One-Step Effective Diffusion Network for Real-World Image Super-ResolutionRongyuan Wu, Lingchen Sun, Zhiyuan Ma, Lei ZhangNeurIPS 2024 · 319 citations
- Diffusion Transformer Meets Multi-Level Wavelet Spectrum for Single Image Super-ResolutionPeng Du, Hui Li, Han Xu, Paul Barom Jeon et al.ICCV 2025 · 3 citations
- Towards Authentic Face Restoration with Iterative Diffusion Models and BeyondYang Zhao, Tingbo Hou, Yu-Chuan Su, Xuhui Jia et al.ICCV 2023 · 30 citations
- Adversarial Diffusion Compression for Real-World Image Super-ResolutionBin Chen, Gehui Li, Rongyuan Wu, Xindong Zhang et al.CVPR 2025
- Visual Autoregressive Modeling for Image Super-ResolutionYunpeng Qu, Kun Yuan, Jinhua Hao, Kai Zhao et al.ICML 2025
