Stepping Forward on the Last Mile
Chen Feng, Jay Zhuo, Parker Zhang, Ramchalam Kinattinkara Ramakrishnan, Zhaocong Yuan, Andrew Zou Li
Abstract
Continuously adapting pre-trained models to local data on resource constrained edge devices is the for model deployment. However, as models increase in size and depth, backpropagation requires a large amount of memory, which becomes prohibitive for edge devices. In addition, most existing low power neural processing engines (e.g., NPUs, DSPs, MCUs, etc.) are designed as fixed-point inference accelerators, without training capabilities. Forward gradients, solely based on directional derivatives computed from two forward calls, have been recently used for model training, with substantial savings in computation and memory. However, the performance of quantized training with fixed-point forward gradients remains unclear. In this paper, we investigate the feasibility of on-device training using fixed-point forward gradients, by conducting comprehensive experiments across a variety of deep learning benchmark tasks in both vision and audio domains. We propose a series of algorithm enhancements that further reduce the memory footprint, and the accuracy gap compared to backpropagation. An empirical study on how training with forward gradients navigates in the loss landscape is further explored. Our results demonstrate that on the last mile of model customization on edge devices, training with fixed-point forward gradients is a feasible and practical approach.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b170f323-308b-40ec-a728-4750e894866dCited by top-tier papers2
- Fine-tuning Quantized Neural Networks with Zeroth-order OptimizationSifeng SHANG, JIAYI ZHOU, Chenyu Lin, Minxian Li et al.ICLR 2026 · 5 citations
- QuZO: Quantized Zeroth-Order Fine-Tuning for Large Language ModelsJiajun Zhou, Yifan Yang, Kai Zhen, Ziyue Liu et al.EMNLP 2025
Builds on10
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Training data-efficient image transformers & distillation through attentionHugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa et al.ICML 2021 · 8,974 citations
- Fine-Tuning Language Models with Just Forward PassesSadhika Malladi, Tianyu Gao, Eshaan Nichani, Alex Damian et al.NeurIPS 2023 · 495 citations
- On-Device Training Under 256KB MemoryJi Lin, Ligeng Zhu, Wei-Ming Chen, Wei-Chen Wang et al.NeurIPS 2022 · 345 citations
- DeepZero: Scaling Up Zeroth-Order Optimization for Deep Model TrainingAochuan Chen, Yimeng Zhang, Jinghan Jia, James Diffenderfer et al.ICLR 2024 · 88 citations
Related papers
- FF-INT8: Efficient Forward-Forward DNN Training on Edge Devices with INT8 PrecisionJingxiao Ma, Priyadarshini Panda, Sherief RedaDAC 2025 · 1 citation
- One-Step Forward and Backtrack: Overcoming Zig-Zagging in Loss-Aware Quantization TrainingLianbo Ma, Yuee Zhou, Jianlun Ma, Guo Yu et al.AAAI 2024 · 5 citations
- Accelerated On-Device Forward Neural Network Training with Module-Wise Descending AsynchronismXiaohan Zhao, Hualin Zhang, Zhouyuan Huo, Bin GuNeurIPS 2023 · 1 citation
- Octo: INT8 Training with Loss-aware Compensation and Backward Quantization for Tiny On-device LearningQihua Zhou, Song Guo, Zhihao Qu, Jingcai Guo et al.USENIX ATC 2021 · 55 citations
- 8-bit Transformer Inference and Fine-tuning for Edge AcceleratorsJeffrey Yu, Kartik Prabhu, Yonatan Urman, Robert M. Radway et al.ASPLOS 2024 · 26 citations
