Perceive, Understand and Restore: Real-World Image Super-Resolution with Autoregressive Multimodal Generative Models
Hongyang Wei, Shuaizheng Liu, Chun Yuan, Lei Zhang
摘要
By leveraging the generative priors from pre-trained text-to-image diffusion models, significant progress has been made in real-world image super-resolution (Real-ISR). However, these methods tend to generate inaccurate and unnatural reconstructions in complex and/or heavily degraded scenes, primarily due to their limited perception and understanding capability of the input low-quality image. To address these limitations, we propose, for the first time to our knowledge, to adapt the pre-trained autoregressive multimodal model such as Lumina-mGPT into a robust Real-ISR model, namely PURE, which Perceives and Understands the input low-quality image, then REstores its high-quality counterpart. Specifically, we implement instruction tuning on Lumina-mGPT to perceive the image degradation level and the relationships between previously generated image tokens and the next token, understand the image content by generating image semantic descriptions, and consequently restore the image by generating high-quality image tokens autoregressively with the collected information. In addition, we reveal that the image token entropy reflects the image structure and present a entropy-based Top-k sampling strategy to optimize the local structure of the image during inference. Experimental results demonstrate that PURE preserves image content while generating realistic details, especially in complex scenes with multiple objects, showcasing the potential of autoregressive multimodal generative models for robust Real-ISR. The model and code will be available at https://github.com/nonwhy/PURE.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper12
- Towards Better & Faster Autoregressive Image Generation: From the Perspective of EntropyXiaoxiao Ma, Feng Zhao, Pengyang Ling, Haibo Qiu 等NeurIPS 2025 · 被引用 12 次
- VQRAE: Representation Quantization Autoencoders for Multimodal Understanding, Generation and ReconstructionSinan Du, Jiahao Guo, Bo Li, Shuhao Cui 等CVPR 2026 · 被引用 11 次
- MICo-150K: A Comprehensive Dataset Advancing Multi-Image CompositionXinyu Wei, Kangrui Cen, Hongyang Wei, Zhen Guo 等CVPR 2026 · 被引用 10 次
- OptMerge: Unifying Multimodal LLM Capabilities and Modalities via Model MergingYongxian Wei, Runxi Cheng, Weike Jin, Enneng Yang 等ICLR 2026 · 被引用 10 次
- UARE: A Unified Vision-Language Model for Image Quality Assessment, Restoration, and EnhancementWeiqi Li, Xuanyu Zhang, Bin Chen, Jingfen Xie 等CVPR 2026 · 被引用 5 次
它引用的顶会 Paper19
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 被引用 11,349 次
- Adding Conditional Control to Text-to-Image Diffusion ModelsLvmin Zhang, Anyi Rao, Maneesh AgrawalaICCV 2023 · 被引用 6,759 次
相关 Paper
- SeeSR: Towards Semantics-Aware Real-World Image Super-ResolutionRongyuan Wu, Tao Yang, Lingchen Sun, Zhengqiang Zhang 等CVPR 2024 · 被引用 119 次
- LD-RPS: Zero-Shot Unified Image Restoration via Latent Diffusion Recurrent Posterior SamplingHuaqiu Li, Yong Wang, Tongwen Huang, Hailang Huang 等ICCV 2025 · 被引用 4 次
- One-Step Effective Diffusion Network for Real-World Image Super-ResolutionRongyuan Wu, Lingchen Sun, Zhiyuan Ma, Lei ZhangNeurIPS 2024 · 被引用 319 次
- BiProLoRA: Bilevel Prompt LoRA for Real Scene RecoveryNan An, Long Ma, Tengyu Ma, Zhu Liu 等CVPR 2026
- Beyond Ground-Truth: Leveraging Image Quality Priors for Real-World Image RestorationFengyang Xiao, Peng Hu, Lei Xu, XingE Guo 等CVPR 2026 · 被引用 5 次
