NIPQ: Noise proxy-based Integrated Pseudo-Quantization
Juncheol Shin, Junhyuk So, Sein Park, Seungyeop Kang, Sungjoo Yoo, Eunhyeok Park
摘要
Straight-through estimator (STE), which enables the gradient flow over the non-differentiable function via approximation, has been favored in studies related to quantization-aware training (QAT). However, STE incurs unstable convergence during QAT, resulting in notable quality degradation in low precision. Recently, pseudoquantization training has been proposed as an alternative approach to updating the learnable parameters using the pseudo-quantization noise instead of STE. In this study, we propose a novel noise proxy-based integrated pseudoquantization (NIPQ) that enables unified support of pseudoquantization for both activation and weight by integrating the idea of truncation on the pseudo-quantization framework. NIPQ updates all of the quantization parameters (e.g., bit-width and truncation boundary) as well as the network parameters via gradient descent without STE instability. According to our extensive experiments, NIPQ outperforms existing quantization algorithms in various vision and language applications by a large margin. * Equal contribution.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper16
- Temporal Dynamic Quantization for Diffusion ModelsJunhyuk So, Jungwon Lee, Daehyun Ahn, Hyungjun Kim 等NeurIPS 2023 · 被引用 109 次
- MetaMix: Meta-State Precision Searcher for Mixed-Precision Activation QuantizationHan-Byul Kim, Joo Hyung Lee, Sungjoo Yoo, Hong-Seok KimAAAI 2024 · 被引用 10 次
- GIFT-SW: Gaussian noise Injected Fine-Tuning of Salient Weights for LLMsMaxim Zhelnin, Viktor Moskvoretskii, Egor Shvetsov, Mariya Krylova 等ACL 2025 · 被引用 7 次
- Improving the Straight-Through Estimator with Zeroth-Order InformationNingfeng Yang, Tor M. AamodtNeurIPS 2025 · 被引用 6 次
- Retraining-free Model Quantization via One-Shot Weight-Coupling LearningChen Tang, Yuan Meng, Jiacheng Jiang, Shuzhao Xie 等CVPR 2024 · 被引用 5 次
它引用的顶会 Paper11
- Learned Step Size quantizationSteven K. Esser, Jeffrey L. McKinstry, Deepika Bablani, Rathinakumar Appuswamy 等ICLR 2020 · 被引用 1,037 次
- HAWQ: Hessian AWare Quantization of Neural Networks With Mixed-PrecisionZhen Dong, Zhewei Yao, Amir Gholami, Michael W. Mahoney 等ICCV 2019 · 被引用 645 次
- HAWQ-V2: Hessian Aware trace-Weighted Quantization of Neural NetworksZhen Dong, Zhewei Yao, Daiyaan Arfeen, Amir Gholami 等NeurIPS 2020 · 被引用 434 次
- HAWQ-V3: Dyadic Neural Network QuantizationZhewei Yao, Zhen Dong, Zhangcheng Zheng, Amir Gholami 等ICML 2021 · 被引用 240 次
- Stabilizing Differentiable Architecture Search via Perturbation-based RegularizationXiangning Chen, Cho-Jui HsiehICML 2020 · 被引用 235 次
相关 Paper
- Network Quantization With Element-Wise Gradient ScalingJunghyup Lee, Dohyung Kim, Bumsub HamCVPR 2021
- Distance-aware QuantizationDohyung Kim, Junghyup Lee, Bumsub HamICCV 2021 · 被引用 42 次
- WinQ: Accelerating Quantization-Aware Training of Language Models Around Saddle PointsDongyue Li, Zechun Liu, Kai Yi, Zhenshuo Zhang 等ICML 2026
- A Statistical Framework for Low-bitwidth Training of Deep Neural NetworksJianfei Chen, Yu Gai, Zhewei Yao, Michael W. Mahoney 等NeurIPS 2020 · 被引用 75 次
- PARQ: Piecewise-Affine Regularized QuantizationLisa Jin, Jianhao Ma, Zechun Liu, Andrey Gromov 等ICML 2025
