Hierarchical Image Tokenization for Multi-Scale Image Super Resolution
Isma Hadji, Enrique Sanchez, Adrian Bulat, Brais Martinez, Georgios Tzimiropoulos
Abstract
We introduce a multi-scale Image Super Resolution (ISR) method building on recent advances in Visual Auto-Regressive (VAR) modeling. VAR models break image tokenization into additive, gradually increasing scales, using Residual Quantization (RQ), an approach that aligns perfectly with our target ISR task. Previous works taking advantage of this synergy suffer from two main shortcomings. First, due to the limitations in RQ, they only generate images at a predefined fixed scale, failing to map intermediate outputs to the corresponding image scales. They also rely on large backbones or a large corpus of annotated data to achieve better performance. To address both shortcomings, we introduce two novel components to the VAR training for ISR, aiming at increasing its flexibility and reducing its complexity. In particular, we introduce a) a Hierarchical Image Tokenization (HIT) approach that progressively represents images at different scales while enforcing token overlap across scales, and b) a Direct Preference Optimization (DPO) regularization term that, relying solely on the (LR,HR) pair, encourages the transformer to produce the latter over the former. Our proposed HIT acts as a strong inductive bias for the VAR training, resulting in a small model (300M params vs 1B params of VARSR), that achieves state-of-the-art results without external training data, and that delivers multi-scale outputs with a single forward pass.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on26
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning et al.NeurIPS 2023 · 10,924 citations
- Adding Conditional Control to Text-to-Image Diffusion ModelsLvmin Zhang, Anyi Rao, Maneesh AgrawalaICCV 2023 · 6,759 citations
- Visual Autoregressive Modeling: Scalable Image Generation via Next-Scale PredictionKeyu Tian, Yi Jiang, Zehuan Yuan, Bingyue Peng et al.NeurIPS 2024 · 1,199 citations
- Designing a Practical Degradation Model for Deep Blind Image Super-ResolutionKai Zhang, Jingyun Liang, Luc Van Gool, Radu TimofteICCV 2021 · 898 citations
Related papers
- DVAR: Dynamic Visual Autoregressive Modeling for Image Super-ResolutionYu Zheng, Kai Zhang, Wei Zhu, Qingguo Liu et al.CVPR 2026
- Visual Autoregressive Modeling for Image Super-ResolutionYunpeng Qu, Kun Yuan, Jinhua Hao, Kai Zhao et al.ICML 2025
- VARestorer: One-Step VAR Distillation for Real-World Image Super-ResolutionYixuan Zhu, Shilin Ma, Haolin Wang, Ao Li et al.ICLR 2026 · 1 citation
- HMAR: Efficient Hierarchical Masked Auto-Regressive Image GenerationHermann Kumbong, Xian Liu, Tsung-Yi Lin, Ming-Yu Liu et al.CVPR 2025
- FastVAR: Linear Visual Autoregressive Modeling Via Cached Token PruningHang Guo, Yawei Li, Taolin Zhang, Jiangshan Wang et al.ICCV 2025 · 5 citations
