Addressing Representation Collapse in Vector Quantized Models with One Linear Layer
Yongxin Zhu, Bocheng Li, Yifei Xin, Zhihua Xia, Linli Xu
摘要
Vector Quantization (VQ) is essential for discretizing continuous representations in unsupervised learning but suffers from representation collapse, causing low codebook utilization and limiting scalability. Existing solutions often rely on complex optimizations or reduce latent dimensionality, which compromises model capacity and fails to fully solve the problem. We identify the root cause as disjoint codebook optimization, where only a few code vectors are updated via gradient descent. To fix this, we propose SimpleVQ, which reparameterizes code vectors through a learnable linear transformation layer over a latent basis, optimizing the entire linear space rather than nearest individual code vectors. Although the multiplication of two linear matrices is equivalent to applying a single linear layer, this simple approach effectively prevents collapse. Extensive experiments on image and audio tasks demonstrate that SimVQ improves codebook usage, is easy to implement, and generalizes well across modalities and architectures. The code is available at https://github.com/youngsheen/SimVQ.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper22
- SemHiTok: A Unified Image Tokenizer via Semantic-Guided Hierarchical Codebook for Multimodal Understanding and GenerationZisheng Chen, Chunwei Wang, Runhui Huang, Hongbin Xu 等ICLR 2026 · 被引用 24 次
- VQRAE: Representation Quantization Autoencoders for Multimodal Understanding, Generation and ReconstructionSinan Du, Jiahao Guo, Bo Li, Shuhao Cui 等CVPR 2026 · 被引用 11 次
- Scalable Training for Vector-Quantized Networks with 100% Codebook UtilizationYifan Chang, Jie Qin, Limeng Qiao, Xiaofeng Wang 等ICLR 2026 · 被引用 10 次
- FACE: A General Framework for Mapping Collaborative Filtering Embeddings into LLM TokensChao Wang, Yixin Song, Jinhui Ye, Chuan Qin 等NeurIPS 2025 · 被引用 7 次
- HiMTok: Learning Hierarchical Mask Tokens for Image Segmentation with Large Multimodal ModelTao Wang, Changxu Cheng, Lingfeng Wang, Senda Chen 等ICCV 2025 · 被引用 5 次
它引用的顶会 Paper20
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- wav2vec 2.0: A Framework for Self-Supervised Learning of Speech RepresentationsAlexei Baevski, Yuhao Zhou, Abdelrahman Mohamed, Michael AuliNeurIPS 2020 · 被引用 9,451 次
- Zero-Shot Text-to-Image GenerationAditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray 等ICML 2021 · 被引用 6,356 次
- Vector-quantized Image Modeling with Improved VQGANJiahui Yu, Xin Li, Jing Yu Koh, Han Zhang 等ICLR 2022 · 被引用 753 次
相关 Paper
- SQ-VAE: Variational Bayes on Discrete Representation with Self-annealed Stochastic QuantizationYuhta Takida, Takashi Shibuya, Wei-Hsiang Liao, Chieh-Hsin Lai 等ICML 2022 · 被引用 99 次
- Unveiling And Addressing Dimensional Collapse In Vector Quantization Models Via Codebook RegularizationFang Zhang, Yongxin Zhu, Yihao Liu, Bin Fu 等ICML 2026
- Online Clustered CodebookChuanxia Zheng, Andrea VedaldiICCV 2023 · 被引用 67 次
- IVQ: Structured and Lightweight Vector Quantization via Binary Hierarchical Composition Inspired by Heda Zuo, Junxian Wu, Fengjie Lu, Pei Chen 等ICML 2026
- Scalable Image Tokenization with Index Backpropagation QuantizationFengyuan Shi, Zhuoyan Luo, Yixiao Ge, Yujiu Yang 等ICCV 2025 · 被引用 7 次
