Improving the Straight-Through Estimator with Zeroth-Order Information
Ningfeng Yang, Tor M. Aamodt
Abstract
We study the problem of training neural networks with quantized parameters. Learning low-precision quantized parameters by enabling computation of gradients via the Straight-Through Estimator (STE) can be challenging. While the STE enables back-propagation, which is a first-order method, recent works have explored the use of zeroth-order (ZO) gradient descent for fine-tuning. We note that the STE provides high-quality biased gradients, and ZO gradients are unbiased but can be expensive. We thus propose First-Order-Guided Zeroth-Order Gradient Descent (FOGZO) that reduces STE bias while reducing computations relative to ZO methods. Empirically, we show FOGZO improves the tradeoff between quality and training time in Quantization-Aware Pre-Training. Specifically, versus STE at the same number of iterations, we show a 1-8% accuracy improvement for DeiT Tiny/Small, 1-2% accuracy improvement on ResNet 18/50, and 1-22 perplexity point improvement for LLaMA models with up to 0.3 billion parameters. For the same loss, FOGZO yields a 796 reduction in computation versus n-SPSA for a 2-layer MLP on MNIST. Code is available at https://github.com/1733116199/fogzo.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 0f67cff5-2641-43c7-8315-efd5157c2269Cited by top-tier papers1
Ask how each one uses itBuilds on24
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Training data-efficient image transformers & distillation through attentionHugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa et al.ICML 2021 · 8,974 citations
- Learned Step Size quantizationSteven K. Esser, Jeffrey L. McKinstry, Deepika Bablani, Rathinakumar Appuswamy et al.ICLR 2020 · 1,037 citations
- Fine-Tuning Language Models with Just Forward PassesSadhika Malladi, Tianyu Gao, Eshaan Nichani, Alex Damian et al.NeurIPS 2023 · 495 citations
- GaLore: Memory-Efficient LLM Training by Gradient Low-Rank ProjectionJiawei Zhao, Zhenyu Zhang, Beidi Chen, Zhangyang Wang et al.ICML 2024 · 433 citations
Related papers
- QuZO: Quantized Zeroth-Order Fine-Tuning for Large Language ModelsJiajun Zhou, Yifan Yang, Kai Zhen, Ziyue Liu et al.EMNLP 2025
- Fine-tuning Quantized Neural Networks with Zeroth-order OptimizationSifeng SHANG, JIAYI ZHOU, Chenyu Lin, Minxian Li et al.ICLR 2026 · 5 citations
- Zeroth-Order Fine-Tuning of LLMs with Transferable Static SparsityWentao Guo, Jikai Long, Yimeng Zeng, Zirui Liu et al.ICLR 2025
- Sparse MeZO: Less Parameters for Better Performance in Zeroth-Order LLM Fine-TuningYong Liu, Zirui Zhu, Chaoyu Gong, Minhao Cheng et al.NeurIPS 2025 · 66 citations
- Robust Training of Neural Networks at Arbitrary Precision and SparsityChengxi Ye, Grace Chu, Yanfeng Liu, Yichi Zhang et al.ICLR 2026 · 2 citations
