BP-Modified Local Loss for Efficient Training of Deep Neural Networks
Lianhai Ren, Qianxiao Li
Abstract
The training of large models is memory-constrained, one direction to relieve this is training using local loss, like GIM, LoCo, and Forward-Forward algorithms. However, the local loss methods often face the issue of slow or non-convergence. In this paper, we propose a novel BP-modified local loss method that uses the true Backward Propagation (BP) gradient to modify the local loss gradient to improve the performance of local loss training. We use the stochastic modified equation to analyze our method and show that modified offset decreases the bias between the BP gradient and local loss gradient, but introduces additional variance, which results in a bias-variance balance. Numerical experiments on full-tuning and LoKr tuning on the ResNet-50 model and LoRA tuning on the ViT-b16 model on CIFAR-100 datasets show 20.5% test top-1 accuracy improvement for the Forward-Forward algorithm, 18.6% improvement for LoCo algorithm and achieve only an average 7.7% of test accuracy loss compared to the BP algorithm, with up to 75% memory savings.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 7df6fdd9-9569-47e0-826c-5481799f798bBuilds on6
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Fine-Tuning Language Models with Just Forward PassesSadhika Malladi, Tianyu Gao, Eshaan Nichani, Alex Damian et al.NeurIPS 2023 · 495 citations
- FedPara: Low-rank Hadamard Product for Communication-Efficient Federated LearningNam Hyeon-Woo, Moon Ye-Bin, Tae-Hyun OhICLR 2022 · 179 citations
- Layer Collaboration in the Forward-Forward AlgorithmGuy Lorberbom, Itai Gat, Yossi Adi, Alexander G. Schwing et al.AAAI 2024 · 22 citations
Related papers
- From Weight-Based to State-Based Fine-Tuning: Further Memory Reduction on LoRA with Parallel ControlChi Zhang, Lianhai Ren, Jingpu Cheng, Qianxiao LiICML 2025
- AdaRankGrad: Adaptive Gradient Rank and Moments for Memory-Efficient LLMs Training and Fine-TuningYehonathan Refael, Jonathan Svirsky, Boris Shustin, Wasim Huleihel et al.ICLR 2025
- AltLoRA: Towards Better Gradient Approximation in Low-Rank Adaptation with Alternating ProjectionsXin Yu, Yujia Wang, Jinghui Chen, Lingzhou XueNeurIPS 2025 · 8 citations
- Thinking Forward: Memory-Efficient Federated Finetuning of Language ModelsKunjal Panchal, Nisarg Parikh, Sunav Choudhary, Lijun Zhang et al.NeurIPS 2024 · 11 citations
- Full Parameter Fine-tuning for Large Language Models with Limited ResourcesKai Lv, Yuqing Yang, Tengxiao Liu, Qipeng Guo et al.ACL 2024 · 61 citations
