One-Step Forward and Backtrack: Overcoming Zig-Zagging in Loss-Aware Quantization Training
Lianbo Ma, Yuee Zhou, Jianlun Ma, Guo Yu, Qing Li
Abstract
Weight quantization is an effective technique to compress deep neural networks for their deployment on edge devices with limited resources. Traditional loss-aware quantization methods commonly use the quantized gradient to replace the full-precision gradient. However, we discover that the gradient error will lead to an unexpected zig-zagging-like issue in the gradient descent learning procedures, where the gradient directions rapidly oscillate or zig-zag, and such issue seriously slows down the model convergence. Accordingly, this paper proposes a one-step forward and backtrack way for loss-aware quantization to get more accurate and stable gradient direction to defy this issue. During the gradient descent learning, a one-step forward search is designed to find the trial gradient of the next-step, which is adopted to adjust the gradient of current step towards the direction of fast convergence. After that, we backtrack the current step to update the full-precision and quantized weights through the current-step gradient and the trial gradient. A series of theoretical analysis and experiments on benchmark deep models have demonstrated the effectiveness and competitiveness of the proposed method, and our method especially outperforms others on the convergence performance.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext a9d4bbdc-1673-4e1e-be59-1de8614c1f51Cited by top-tier papers2
- AIM: Software and Hardware Co-design for Architecture-level IR-drop Mitigation in High-performance PIMYuanpeng Zhang, Xing Hu, Xi Chen, Zhihang Yuan et al.ISCA 2025 · 1 citation
- No Retraining at Edge: Efficient Resource-Aware Mixed-Precision Quantization via Federated Supernet LearningLianbo Ma, Yonghui Su, Nan Li, Xingwei WangICML 2026
Builds on7
- Differentiable Soft Quantization: Bridging Full-Precision and Low-Bit Neural NetworksRuihao Gong, Xianglong Liu, Shenghu Jiang, Tianxiang Li et al.ICCV 2019 · 540 citations
- Overcoming Oscillations in Quantization-Aware TrainingMarkus Nagel, Marios Fournarakis, Yelysei Bondarenko, Tijmen BlankevoortICML 2022 · 163 citations
- Don't Waste Your Bits! Squeeze Activations and Gradients for Deep Neural Networks via TinyScriptFangcheng Fu, Yuzheng Hu, Yihan He, Jiawei Jiang et al.ICML 2020 · 73 citations
- Progressive Network Grafting for Few-Shot Knowledge DistillationChengchao Shen, Xinchao Wang, Youtan Yin, Jie Song et al.AAAI 2021 · 55 citations
- Optimal Clipping and Magnitude-aware Differentiation for Improved Quantization-aware TrainingCharbel Sakr, Steve Dai, Rangharajan Venkatesan, Brian Zimmer et al.ICML 2022 · 53 citations
Related papers
- Stepping Forward on the Last MileChen Feng, Jay Zhuo, Parker Zhang, Ramchalam Kinattinkara Ramakrishnan et al.NeurIPS 2024 · 4 citations
- Quantized Compressive Sampling of Stochastic Gradients for Efficient Communication in Distributed Deep LearningAfshin Abdi, Faramarz FekriAAAI 2020 · 32 citations
- Adaptive Loss-Aware Quantization for Multi-Bit NetworksZhongnan Qu, Zimu Zhou, Yun Cheng, Lothar ThieleCVPR 2020
- FF-INT8: Efficient Forward-Forward DNN Training on Edge Devices with INT8 PrecisionJingxiao Ma, Priyadarshini Panda, Sherief RedaDAC 2025 · 1 citation
- WeightGrad: Geo-Distributed Data Analysis Using Quantization for Faster Convergence and Better AccuracySyeda Nahida Akter, Muhammad Abdullah AdnanKDD 2020 · 8 citations
