CODA: Repurposing Continuous VAEs for Discrete Tokenization
Zeyu Liu, Zanlin Ni, Yeguo Hua, Xin Deng, Xiao Ma, Cheng Zhong, Gao Huang
摘要
Discrete visual tokenizers transform images into a sequence of tokens, enabling token-based visual generation akin to language models. However, this process is inherently challenging, as it requires both compressing visual signals into a compact representation and discretizing them into a fixed set of codes. Traditional discrete tokenizers typically learn the two tasks jointly, often leading to unstable training, low codebook utilization, and limited reconstruction quality. In this paper, we introduce CODA (CO ntinuous-toDiscrete Adaptation), a framework that decouples compression and discretization. Instead of training discrete tokenizers from scratch, CODA adapts off-the-shelf continuous VAEs—already optimized for perceptual compression—into discrete tokenizers via a carefully designed discretization process. By primarily focusing on discretization, CODA ensures stable and efficient training while retaining the strong visual fidelity of continuous VAEs. Empirically, with less training budget than standard VQGAN, our approach achieves a remarkable codebook utilization of and notable reconstruction FID (rFID) of and for and compression on ImageNet benchmark.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Linear Differential Vision Transformer: Learning Visual Contrasts via Pairwise DifferentialsYifan Pu, Jixuan Ying, Qixiu Li, Tianzhu Ye 等NeurIPS 2025 · 被引用 9 次
- Soft-Di[M]O: Improving One-Step Discrete Image Generation with Soft EmbeddingsYuanzhi Zhu, Xi Wang, Stéphane Lathuilière, Vicky KalogeitonICLR 2026 · 被引用 4 次
- IMG: Calibrating Diffusion Models via Implicit Multimodal GuidanceJiayi Guo, Chuanhao Yan, Xingqian Xu, Yulin Wang 等ICCV 2025 · 被引用 4 次
它引用的顶会 Paper34
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 被引用 13,211 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 被引用 11,349 次
相关 Paper
- Bridging Continuous and Discrete Tokens for Autoregressive Visual GenerationYuqing Wang, Zhijie Lin, Yao Teng, Yuanzhi Zhu 等ICCV 2025 · 被引用 1 次
- VQ-Transplant: Efficient VQ-Module Integration for Pre-trained Visual TokenizersXianghong Fang, Yuan Yuan, Dehan Kong, Tim G. J. RudnerICLR 2026 · 被引用 2 次
- WeTok: Powerful Discrete Tokenization for High-Fidelity Visual ReconstructionShaobin Zhuang, Yiwei Guo, Fangyikang Wang, Canmiao Fu 等ICLR 2026 · 被引用 9 次
- SoftVQ-VAE: Efficient 1-Dimensional Continuous TokenizerHao Chen, Ze Wang, Xiang Li, Ximeng Sun 等CVPR 2025
- Autoregressive Image Generation with Masked Bit ModelingQihang Yu, Qihao Liu, Ju He, Xinyang Zhang 等ICML 2026 · 被引用 5 次
