VQ-Transplant: Efficient VQ-Module Integration for Pre-trained Visual Tokenizers
Xianghong Fang, Yuan Yuan, Dehan Kong, Tim G. J. Rudner
Abstract
Vector Quantization (VQ) underpins modern discrete visual tokenization. However, training quantization modules for state-of-the-art VQ-based models requires significant computational resources which, in practice, all but prevents the development of novel, cutting-edge VQ techniques under resource constraints. To address this limitation, we propose VQ-Transplant, a simple framework that enables plug-and-play integration of new VQ modules into frozen, pre-trained tokenizers by replacing their native VQ modules. Crucially, the proposed transplantation process preserves all encoder-decoder parameters, obviating the need for costly end-to-end retraining when modifying the quantization method. To mitigate decoder-quantization mismatch, we introduce a lightweight decoder adaptation strategy (trained for only 5 epochs on ImageNet-1k) to align feature priors with the new quantization space. In our empirical evaluation, we find that VQ-Transplant allows obtaining near state-of-the-art reconstruction fidelity for industry-level models like VAR while reducing the training cost by 95%. VQ-Transplant democratizes quantization research by enabling resource-efficient integration of novel VQ techniques while matching industry-level reconstruction performance. Code and models are available at VQ-Transplant. INTRODUCTION Vector Quantization (VQ) is a cornerstone of modern discrete visual tokenization frameworks, enabling efficient learning of discrete representations critical for downstream tasks including visual generation (van den Oord et al.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on31
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 11,743 citations
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou et al.ICCV 2021 · 8,921 citations
- Zero-Shot Text-to-Image GenerationAditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray et al.ICML 2021 · 6,356 citations
Related papers
- CODA: Repurposing Continuous VAEs for Discrete TokenizationZeyu Liu, Zanlin Ni, Yeguo Hua, Xin Deng et al.ICCV 2025 · 9 citations
- Scalable Image Tokenization with Index Backpropagation QuantizationFengyuan Shi, Zhuoyan Luo, Yixiao Ge, Yujiu Yang et al.ICCV 2025 · 7 citations
- Bridging Continuous and Discrete Tokens for Autoregressive Visual GenerationYuqing Wang, Zhijie Lin, Yao Teng, Yuanzhi Zhu et al.ICCV 2025 · 1 citation
- VQRAE: Representation Quantization Autoencoders for Multimodal Understanding, Generation and ReconstructionSinan Du, Jiahao Guo, Bo Li, Shuhao Cui et al.CVPR 2026 · 11 citations
- Scalable Training for Vector-Quantized Networks with 100% Codebook UtilizationYifan Chang, Jie Qin, Limeng Qiao, Xiaofeng Wang et al.ICLR 2026 · 10 citations
