SQ-VAE: Variational Bayes on Discrete Representation with Self-annealed Stochastic Quantization
Yuhta Takida, Takashi Shibuya, Wei-Hsiang Liao, Chieh-Hsin Lai, Junki Ohmura, Toshimitsu Uesaka, Naoki Murata, Shusuke Takahashi, Toshiyuki Kumakura, Yuki Mitsufuji
Abstract
One noted issue of vector-quantized variational autoencoder (VQ-VAE) is that the learned discrete representation uses only a fraction of the full capacity of the codebook, also known as codebook collapse. We hypothesize that the training scheme of VQ-VAE, which involves some carefully designed heuristics, underlies this issue. In this paper, we propose a new training scheme that extends the standard VAE via novel stochastic dequantization and quantization, called stochastically quantized variational autoencoder (SQ-VAE). In SQ-VAE, we observe a trend that the quantization is stochastic at the initial stage of the training but gradually converges toward a deterministic quantization, which we call self-annealing. Our experiments show that SQ-VAE improves codebook utilization without using common heuristics. Furthermore, we empirically show that SQ-VAE is superior to VAE and VQ-VAE in vision- and speech-related tasks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext ef5dcb4d-4fef-4863-9adb-240b89ed869aCited by top-tier papers34
- Finite Scalar Quantization: VQ-VAE Made SimpleFabian Mentzer, David Minnen, Eirikur Agustsson, Michael TschannenICLR 2024 · 442 citations
- Straightening Out the Straight-Through Estimator: Overcoming Optimization Challenges in Vector Quantized NetworksMinyoung Huh, Brian Cheung, Pulkit Agrawal, Phillip IsolaICML 2023 · 104 citations
- Mega-TTS 2: Boosting Prompting Mechanisms for Zero-Shot Speech SynthesisZiyue Jiang, Jinglin Liu, Yi Ren, Jinzheng He et al.ICLR 2024 · 75 citations
- Online Clustered CodebookChuanxia Zheng, Andrea VedaldiICCV 2023 · 67 citations
- SAN: Inducing Metrizability of GAN with Discriminative Normalized Linear LayerYuhta Takida, Masaaki Imaizumi, Takashi Shibuya, Chieh-Hsin Lai et al.ICLR 2024 · 28 citations
Builds on4
- Zero-Shot Text-to-Image GenerationAditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray et al.ICML 2021 · 6,356 citations
- From Variational to Deterministic AutoencodersPartha Ghosh, Mehdi S. M. Sajjadi, Antonio Vergari, Michael J. Black et al.ICLR 2020 · 298 citations
- Vector Quantization-Based Regularization for AutoencodersHanwei Wu, Markus FlierlAAAI 2020 · 33 citations
- Taming Transformers for High-Resolution Image SynthesisPatrick Esser, Robin Rombach, Björn OmmerCVPR 2021
Related papers
- Restructuring Vector Quantization with the Rotation TrickChristopher Fifty, Ronald Guenther Junkins, Dennis Duan, Aniketh Iyengar et al.ICLR 2025 · 1 citation
- Addressing Representation Collapse in Vector Quantized Models with One Linear LayerYongxin Zhu, Bocheng Li, Yifei Xin, Zhihua Xia et al.ICCV 2025 · 5 citations
- Scalable Image Tokenization with Index Backpropagation QuantizationFengyuan Shi, Zhuoyan Luo, Yixiao Ge, Yujiu Yang et al.ICCV 2025 · 7 citations
- Hierarchical Vector Quantized Graph Autoencoder with Annealing-Based Code SelectionLong Zeng, Jianxiang Yu, Jiapeng Zhu, Qingsong Zhong et al.WWW 2025 · 8 citations
- Hierarchical Quantized AutoencodersWill Williams, Sam Ringer, Tom Ash, David MacLeod et al.NeurIPS 2020 · 90 citations
