Finite Scalar Quantization: VQ-VAE Made Simple
Fabian Mentzer, David Minnen, Eirikur Agustsson, Michael Tschannen
摘要
We propose to replace vector quantization (VQ) in the latent representation of VQ-VAEs with a simple scheme termed finite scalar quantization (FSQ), where we project the VAE representation down to a few dimensions (typically less than 10). Each dimension is quantized to a small set of fixed values, leading to an (implicit) codebook given by the product of these sets. By appropriately choosing the number of dimensions and values each dimension can take, we obtain the same codebook size as in VQ. On top of such discrete representations, we can train the same models that have been trained on VQ-VAE representations. For example, autoregressive and masked transformer models for image generation, multimodal generation, and dense prediction computer vision tasks. Concretely, we employ FSQ with MaskGIT for image generation, and with UViM for depth estimation, colorization, and panoptic segmentation. Despite the much simpler design of FSQ, we obtain competitive performance in all these tasks. We emphasize that FSQ does not suffer from codebook collapse and does not need the complex machinery employed in VQ (commitment losses, codebook reseeding, code splitting, entropy penalties, etc.) to learn expressive discrete representations. Code on GitHub.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper189
- Visual Autoregressive Modeling: Scalable Image Generation via Next-Scale PredictionKeyu Tian, Yi Jiang, Zehuan Yuan, Bingyue Peng 等NeurIPS 2024 · 被引用 1,199 次
- Autoregressive Image Generation without Vector QuantizationTianhong Li, Yonglong Tian, He Li, Mingyang Deng 等NeurIPS 2024 · 被引用 758 次
- An Image is Worth 32 Tokens for Reconstruction and GenerationQihang Yu, Mark Weber, Xueqing Deng, Xiaohui Shen 等NeurIPS 2024 · 被引用 331 次
- Diffusion Transformers with Representation AutoencodersBoyang Zheng, Nanye Ma, Shengbang Tong, Saining XieICLR 2026 · 被引用 288 次
- SMART: Scalable Multi-agent Real-time Motion Generation via Next-token PredictionWei Wu, Xiaoxin Feng, Ziyan Gao, Yuheng KanNeurIPS 2024 · 被引用 104 次
它引用的顶会 Paper18
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 被引用 13,211 次
- BEiT: BERT Pre-Training of Image TransformersHangbo Bao, Li Dong, Songhao Piao, Furu WeiICLR 2022 · 被引用 3,632 次
- Vector-quantized Image Modeling with Improved VQGANJiahui Yu, Xin Li, Jing Yu Koh, Han Zhang 等ICLR 2022 · 被引用 753 次
- vq-wav2vec: Self-Supervised Learning of Discrete Speech RepresentationsAlexei Baevski, Steffen Schneider, Michael AuliICLR 2020 · 被引用 730 次
- High-Fidelity Generative Image CompressionFabian Mentzer, George Toderici, Michael Tschannen, Eirikur AgustssonNeurIPS 2020 · 被引用 675 次
相关 Paper
- MoVQ: Modulating Quantized Vectors for High-Fidelity Image GenerationChuanxia Zheng, Tung-Long Vuong, Jianfei Cai, Dinh PhungNeurIPS 2022 · 被引用 156 次
- SQ-VAE: Variational Bayes on Discrete Representation with Self-annealed Stochastic QuantizationYuhta Takida, Takashi Shibuya, Wei-Hsiang Liao, Chieh-Hsin Lai 等ICML 2022 · 被引用 99 次
- Unveiling And Addressing Dimensional Collapse In Vector Quantization Models Via Codebook RegularizationFang Zhang, Yongxin Zhu, Yihao Liu, Bin Fu 等ICML 2026
- Learning to Quantize for Training Vector-Quantized NetworksPeijia Qin, Jianguo ZhangICML 2025
- Online Clustered CodebookChuanxia Zheng, Andrea VedaldiICCV 2023 · 被引用 67 次
