Lune

ICCV2025Top-tier venue

CODA: Repurposing Continuous VAEs for Discrete Tokenization

Zeyu Liu, Zanlin Ni, Yeguo Hua, Xin Deng, Xiao Ma, Cheng Zhong, Gao Huang

2025Year
9Citations
3Top-tier citations

Abstract

Discrete visual tokenizers transform images into a sequence of tokens, enabling token-based visual generation akin to language models. However, this process is inherently challenging, as it requires both compressing visual signals into a compact representation and discretizing them into a fixed set of codes. Traditional discrete tokenizers typically learn the two tasks jointly, often leading to unstable training, low codebook utilization, and limited reconstruction quality. In this paper, we introduce CODA (CO ntinuous-toDiscrete Adaptation), a framework that decouples compression and discretization. Instead of training discrete tokenizers from scratch, CODA adapts off-the-shelf continuous VAEs—already optimized for perceptual compression—into discrete tokenizers via a carefully designed discretization process. By primarily focusing on discretization, CODA ensures stable and efficient training while retaining the strong visual fidelity of continuous VAEs. Empirically, with 6×\mathbf{6} \times less training budget than standard VQGAN, our approach achieves a remarkable codebook utilization of 100%\mathbf{1 0 0 \%} and notable reconstruction FID (rFID) of 0.43\mathbf{0. 4 3} and 1.34\mathbf{1. 3 4} for 8×8 \times and 16×16 \times compression on ImageNet 256×256256 \times 256 benchmark.

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

lune papers fulltext ae82087b-e94e-42fb-b903-c0b2cc2b3174

Cited by top-tier papers3

Ask how each one uses it

Builds on34

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines