Chain-of-Zoom: Extreme Super-Resolution via Scale Autoregression and Preference Alignment
Bryan Sangwoo Kim, Jeongsol Kim, Jong Chul Ye
Abstract
Modern single-image super-resolution (SISR) models deliver photo-realistic results at the scale factors on which they are trained, but collapse when asked to magnify far beyond that regime. We address this scalability bottleneck with Chain-of-Zoom (CoZ), a model-agnostic framework that factorizes SISR into an autoregressive chain of intermediate scale-states with multi-scale-aware prompts. CoZ repeatedly re-uses a backbone SR model, decomposing the conditional probability into tractable sub-problems to achieve extreme resolutions without additional training. Because visual cues diminish at high magnifications, we augment each zoom step with multi-scale-aware text prompts generated by a vision-language model (VLM). The prompt extractor itself is fine-tuned using Generalized Reward Policy Optimization (GRPO) with a critic VLM, aligning text guidance towards human preference. Experiments show that a standard 4x diffusion SR model wrapped in CoZ attains beyond 256x enlargement with high perceptual quality and fidelity. Project Page: https://bryanswkim.github.io/chain-of-zoom/.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 5c2f5ab3-ec20-4f0f-926b-be66729e3b3cCited by top-tier papers2
- WonderZoom: Multi-Scale 3D World GenerationJin Cao, Hong-Xing Yu, Jiajun WuCVPR 2026 · 1 citation
- GaussianZoom: Progressive Zoom-in Generative 3D Gaussian Splatting with Geometric and Semantic GuidanceJiale Shi, Jiarui Hu, Zesong Yang, Kaixuan Luan et al.CVPR 2026
Builds on20
- Scaling Rectified Flow Transformers for High-Resolution Image SynthesisPatrick Esser, Sumith Kulal, Andreas Blattmann, Rahim Entezari et al.ICML 2024 · 3,620 citations
- MUSIQ: Multi-scale Image Quality TransformerJunjie Ke, Qifei Wang, Yilin Wang, Peyman Milanfar et al.ICCV 2021 · 1,325 citations
- Exploring CLIP for Assessing the Look and Feel of ImagesJianyi Wang, Kelvin C. K. Chan, Chen Change LoyAAAI 2023 · 1,208 citations
- Visual Autoregressive Modeling: Scalable Image Generation via Next-Scale PredictionKeyu Tian, Yi Jiang, Zehuan Yuan, Bingyue Peng et al.NeurIPS 2024 · 1,199 citations
- Improving Diffusion Models for Inverse Problems using Manifold ConstraintsHyungjin Chung, Byeongsu Sim, Dohoon Ryu, Jong Chul YeNeurIPS 2022 · 738 citations
Related papers
- Generative Powers of TenXiaojuan Wang, Janne Kontkanen, Brian Curless, Steven M. Seitz et al.CVPR 2024 · 3 citations
- LENS: Learning to Segment Anything with Unified Reinforced ReasoningLianghui Zhu, Bin Ouyang, Yuxuan Zhang, Tianheng Cheng et al.AAAI 2026 · 7 citations
- HiCoGen: Hierarchical Compositional Text-to-Image Generation in Diffusion Models via Reinforcement LearningHongji Yang, Yucheng Zhou, Wencheng Han, Runzhou Tao et al.CVPR 2026 · 4 citations
- QG-CoC: Question-Guided Chain-of-Captions for Large Multimodal ModelsKuei-Chun Kao, Hsu Tzu-Yin, Yunqi Hong, Ruochen Wang et al.EMNLP 2025
- CASR: A Robust Cyclic Framework for Arbitrary Large-Scale Super-Resolution with Distribution Alignment and Self-Similarity AwarenessWenhao Guo, Zhaoran Zhao, Peng Lu, Sheng Li et al.CVPR 2026 · 1 citation
