SQ-VAE: Variational Bayes on Discrete Representation with Self-annealed Stochastic Quantization
Yuhta Takida, Takashi Shibuya, Wei-Hsiang Liao, Chieh-Hsin Lai, Junki Ohmura, Toshimitsu Uesaka, Naoki Murata, Shusuke Takahashi, Toshiyuki Kumakura, Yuki Mitsufuji
摘要
One noted issue of vector-quantized variational autoencoder (VQ-VAE) is that the learned discrete representation uses only a fraction of the full capacity of the codebook, also known as codebook collapse. We hypothesize that the training scheme of VQ-VAE, which involves some carefully designed heuristics, underlies this issue. In this paper, we propose a new training scheme that extends the standard VAE via novel stochastic dequantization and quantization, called stochastically quantized variational autoencoder (SQ-VAE). In SQ-VAE, we observe a trend that the quantization is stochastic at the initial stage of the training but gradually converges toward a deterministic quantization, which we call self-annealing. Our experiments show that SQ-VAE improves codebook utilization without using common heuristics. Furthermore, we empirically show that SQ-VAE is superior to VAE and VQ-VAE in vision- and speech-related tasks.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper34
- Finite Scalar Quantization: VQ-VAE Made SimpleFabian Mentzer, David Minnen, Eirikur Agustsson, Michael TschannenICLR 2024 · 被引用 442 次
- Straightening Out the Straight-Through Estimator: Overcoming Optimization Challenges in Vector Quantized NetworksMinyoung Huh, Brian Cheung, Pulkit Agrawal, Phillip IsolaICML 2023 · 被引用 104 次
- Mega-TTS 2: Boosting Prompting Mechanisms for Zero-Shot Speech SynthesisZiyue Jiang, Jinglin Liu, Yi Ren, Jinzheng He 等ICLR 2024 · 被引用 75 次
- Online Clustered CodebookChuanxia Zheng, Andrea VedaldiICCV 2023 · 被引用 67 次
- SAN: Inducing Metrizability of GAN with Discriminative Normalized Linear LayerYuhta Takida, Masaaki Imaizumi, Takashi Shibuya, Chieh-Hsin Lai 等ICLR 2024 · 被引用 28 次
它引用的顶会 Paper4
- Zero-Shot Text-to-Image GenerationAditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray 等ICML 2021 · 被引用 6,356 次
- From Variational to Deterministic AutoencodersPartha Ghosh, Mehdi S. M. Sajjadi, Antonio Vergari, Michael J. Black 等ICLR 2020 · 被引用 298 次
- Vector Quantization-Based Regularization for AutoencodersHanwei Wu, Markus FlierlAAAI 2020 · 被引用 33 次
- Taming Transformers for High-Resolution Image SynthesisPatrick Esser, Robin Rombach, Björn OmmerCVPR 2021
相关 Paper
- Restructuring Vector Quantization with the Rotation TrickChristopher Fifty, Ronald Guenther Junkins, Dennis Duan, Aniketh Iyengar 等ICLR 2025 · 被引用 1 次
- Addressing Representation Collapse in Vector Quantized Models with One Linear LayerYongxin Zhu, Bocheng Li, Yifei Xin, Zhihua Xia 等ICCV 2025 · 被引用 5 次
- Scalable Image Tokenization with Index Backpropagation QuantizationFengyuan Shi, Zhuoyan Luo, Yixiao Ge, Yujiu Yang 等ICCV 2025 · 被引用 7 次
- Hierarchical Vector Quantized Graph Autoencoder with Annealing-Based Code SelectionLong Zeng, Jianxiang Yu, Jiapeng Zhu, Qingsong Zhong 等WWW 2025 · 被引用 8 次
- Hierarchical Quantized AutoencodersWill Williams, Sam Ringer, Tom Ash, David MacLeod 等NeurIPS 2020 · 被引用 90 次
