Visual Autoregressive Modeling for Image Super-Resolution
Yunpeng Qu, Kun Yuan, Jinhua Hao, Kai Zhao, Qizhi Xie, Ming Sun, Chao Zhou
Abstract
Image Super-Resolution (ISR) has seen significant progress with the introduction of remarkable generative models. However, challenges such as the trade-off issues between fidelity and realism, as well as computational complexity, have also posed limitations on their application. Building upon the tremendous success of autoregressive models in the language domain, we propose VARSR, a novel visual autoregressive modeling for ISR framework with the form of next-scale prediction. To effectively integrate and preserve semantic information in low-resolution images, we propose using prefix tokens to incorporate the condition. Scale-aligned Rotary Positional Encodings are introduced to capture spatial structures, and the Diffusion Refiner is utilized for modeling quantization residual loss to achieve pixel-level fidelity. Image-based Classifier-free Guidance is proposed to guide the generation of more realistic images. Furthermore, we collect large-scale data and design a training process to obtain robust generative priors. Quantitative and qualitative results show that VARSR is capable of generating high-fidelity and high-realism images with more efficiency than diffusion-based methods. Our codes are released at https: //github.com/quyp2000/VARSR .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 7d972ae0-e3df-46f9-bab5-99e375d36a79Cited by top-tier papers16
- RestoreVAR: Visual Autoregressive Generation for All-in-One Image RestorationSudarshan Rajagopalan, Kartik Narayan, Vishal M. PatelICLR 2026 · 17 citations
- SCALAR: Scale-wise Controllable Visual Autoregressive LearningRyan Xu, Dongyang Jin, Yancheng Bai, Rui Lan et al.AAAI 2026 · 16 citations
- Semantic Context Matters: Improving Conditioning for Autoregressive ModelsDongyang Jin, Ryan Xu, Jianhao Zeng, Rui Lan et al.CVPR 2026 · 12 citations
- Progressive Supernet Training for Efficient Visual Autoregressive ModelingXiaoyue Chen, Yuling Shi, Kaiyuan Li, Huandong Wang et al.CVPR 2026 · 7 citations
- Toward Real-world Infrared Image Super-Resolution: A Unified Autoregressive Framework and Benchmark DatasetYang Zou, Jun Ma, Zhidong Jiao, Xingyuan Li et al.CVPR 2026 · 4 citations
Builds on23
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Adding Conditional Control to Text-to-Image Diffusion ModelsLvmin Zhang, Anyi Rao, Maneesh AgrawalaICCV 2023 · 6,759 citations
- Scalable Diffusion Models with TransformersWilliam Peebles, Saining XieICCV 2023 · 5,568 citations
Related papers
- DVAR: Dynamic Visual Autoregressive Modeling for Image Super-ResolutionYu Zheng, Kai Zhang, Wei Zhu, Qingguo Liu et al.CVPR 2026
- VOSR: A Vision-Only Generative Model for Image Super-ResolutionRongyuan Wu, Lingchen Sun, Zhengqiang Zhang, Xiangtao Kong et al.CVPR 2026 · 3 citations
- SSG: Scaled Spatial Guidance for Multi-Scale Visual Autoregressive GenerationYoungwoo Shin, Jiwan Hur, Junmo KimICLR 2026 · 1 citation
- Hierarchical Image Tokenization for Multi-Scale Image Super ResolutionIsma Hadji, Enrique Sanchez, Adrian Bulat, Brais Martinez et al.ICML 2026
- Autoregressive-based Progressive Coding for Ultra-Low Bitrate Image CompressionZiyuan Zhang, Yichong Xia, Bin Chen, Tianwei Zhang et al.ICLR 2026
