Straightening Out the Straight-Through Estimator: Overcoming Optimization Challenges in Vector Quantized Networks
Minyoung Huh, Brian Cheung, Pulkit Agrawal, Phillip Isola
摘要
This work examines the challenges of training neural networks using vector quantization using straight-through estimation. We find that a primary cause of training instability is the discrepancy between the model embedding and the codevector distribution. We identify the factors that contribute to this issue, including the codebook gradient sparsity and the asymmetric nature of the commitment loss, which leads to misaligned codevector assignments. We propose to address this issue via affine re-parameterization of the code vectors. Additionally, we introduce an alternating optimization to reduce the gradient error introduced by the straight-through estimation. Moreover, we propose an improvement to the commitment loss to ensure better alignment between the codebook representation and the model embedding. These optimization methods improve the mathematical approximation of the straightthrough estimation and, ultimately, the model performance. We demonstrate the effectiveness of our methods on several common model architectures, such as AlexNet, ResNet, and ViT, across various tasks, including image classification and generative modeling. Project page: minyoungg.github.io/vqtorch * Equal contribution 1 MIT CSAIL 2 MIT BCS. Correspondence to: Minyoung Huh
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper39
- Finite Scalar Quantization: VQ-VAE Made SimpleFabian Mentzer, David Minnen, Eirikur Agustsson, Michael TschannenICLR 2024 · 被引用 442 次
- GFT: Graph Foundation Model with Transferable Tree VocabularyZehong Wang, Zheyuan Zhang, Nitesh V. Chawla, Chuxu Zhang 等NeurIPS 2024 · 被引用 108 次
- Copilot4D: Learning Unsupervised World Models for Autonomous Driving via Discrete DiffusionLunjun Zhang, Yuwen Xiong, Ze Yang, Sergio Casas 等ICLR 2024 · 被引用 105 次
- Learning to Act without ActionsDominik Schmidt, Minqi JiangICLR 2024 · 被引用 98 次
- UniAudio 1.5: Large Language Model-Driven Audio Codec is A Few-Shot Audio Task LearnerDongchao Yang, Haohan Guo, Yuanyuan Wang, Rongjie Huang 等NeurIPS 2024 · 被引用 55 次
它引用的顶会 Paper11
- Zero-Shot Text-to-Image GenerationAditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray 等ICML 2021 · 被引用 6,356 次
- Unsupervised Learning of Visual Features by Contrasting Cluster AssignmentsMathilde Caron, Ishan Misra, Julien Mairal, Priya Goyal 等NeurIPS 2020 · 被引用 5,249 次
- Self-labelling via simultaneous clustering and representation learningYuki Markus Asano, Christian Rupprecht, Andrea VedaldiICLR 2020 · 被引用 873 次
- Vector-quantized Image Modeling with Improved VQGANJiahui Yu, Xin Li, Jing Yu Koh, Han Zhang 等ICLR 2022 · 被引用 753 次
- When Vision Transformers Outperform ResNets without Pre-training or Strong Data AugmentationsXiangning Chen, Cho-Jui Hsieh, Boqing GongICLR 2022 · 被引用 388 次
相关 Paper
- Correcting Quantization-Induced Gradient Mismatch in Neural Image CompressionChanghao Peng, Yuqi Ye, Wei GaoAAAI 2026
- Restructuring Vector Quantization with the Rotation TrickChristopher Fifty, Ronald Guenther Junkins, Dennis Duan, Aniketh Iyengar 等ICLR 2025 · 被引用 1 次
- DiVeQ: Differentiable Vector Quantization Using the Reparameterization TrickMohammad Hassan Vali, Tom Bäckström, Arno SolinICLR 2026 · 被引用 6 次
- Oscillation-free Quantization for Low-bit Vision TransformersShih-Yang Liu, Zechun Liu, Kwang-Ting ChengICML 2023 · 被引用 63 次
- VAEVQ: Enhancing Discrete Visual Tokenization Through Variational ModelingSicheng Yang, Xing Hu, Qiang Wu, Dawei YangAAAI 2026
