Texture Vector-Quantization and Reconstruction Aware Prediction for Generative Super-Resolution
Qifan Li, Jiale Zou, Jinhua Zhang, Wei Long, Xingyu Zhou, Shuhang Gu
Abstract
Vector-quantized based models have recently demonstrated strong potential for visual prior modeling. However, existing VQ-based methods simply encode visual features with nearest codebook items and train index predictor with code-level supervision. Due to the richness of visual signal, VQ encoding often leads to large quantization error. Furthermore, training predictor with code-level supervision can not take the final reconstruction errors into consideration, result in sub-optimal prior modeling accuracy. In this paper we address the above two issues and propose a Texture Vector-Quantization and a Reconstruction Aware Prediction strategy. The texture vector-quantization strategy leverages the task character of superresolution and only introduce codebook to model the prior of missing textures. While the reconstruction aware prediction strategy makes use of the straightthrough estimator to directly train index predictor with image-level supervision. Our proposed generative SR model (TVQ&RAP) is able to deliver photo-realistic SR results with small computational cost. di se nt an gl in g Vanilla Codebook: Modeling a complex feature space containing both structures and textures. encoding F e a tu r e (a) Vanilla Vector Quantization Lookup Lookup encoding F e a tu r e Structures inherently in LR Texture Codebook: Modeling a simple feature space by remove structures inherently in LR. (b) Texture Vector Quantization
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 1394a022-6a0d-4aab-8ef0-7fa5ae0e90c0Cited by top-tier papers1
Ask how each one uses itBuilds on22
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Zero-Shot Text-to-Image GenerationAditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray et al.ICML 2021 · 6,356 citations
- MUSIQ: Multi-scale Image Quality TransformerJunjie Ke, Qifei Wang, Yilin Wang, Peyman Milanfar et al.ICCV 2021 · 1,325 citations
- Exploring CLIP for Assessing the Look and Feel of ImagesJianyi Wang, Kelvin C. K. Chan, Chen Change LoyAAAI 2023 · 1,208 citations
Related papers
- LG-VQ: Language-Guided Codebook LearningGuotao Liang, Baoquan Zhang, Yaowei Wang, Yunming Ye et al.NeurIPS 2024 · 14 citations
- Scalable Image Tokenization with Index Backpropagation QuantizationFengyuan Shi, Zhuoyan Luo, Yixiao Ge, Yujiu Yang et al.ICCV 2025 · 7 citations
- Visual Autoregressive Modeling for Image Super-ResolutionYunpeng Qu, Kun Yuan, Jinhua Hao, Kai Zhao et al.ICML 2025
- VAEVQ: Enhancing Discrete Visual Tokenization Through Variational ModelingSicheng Yang, Xing Hu, Qiang Wu, Dawei YangAAAI 2026
- VQ-Transplant: Efficient VQ-Module Integration for Pre-trained Visual TokenizersXianghong Fang, Yuan Yuan, Dehan Kong, Tim G. J. RudnerICLR 2026 · 2 citations
