AdaBet: Gradient-free Layer Selection for Efficient Training of Deep Neural Networks
Irene Tenison, Soumyajit Chatterjee, Fahim Kawsar, Mohammad Malekzadeh
Abstract
To utilize pre-trained neural networks on edge and mobile devices, we often require efficient adaptation to userspecific runtime data distributions while operating under limited compute and memory resources. On-device retraining with a target dataset can facilitate such adaptations; however, it remains impractical due to the increasing depth of modern neural nets, as well as the computational overhead associated with gradient-based optimization across all layers. Current approaches reduce training cost by selecting a subset of layers for retraining; however, they rely on labeled data, at least one full-model backpropagation, or server-side meta-training, limiting their suitability for constrained devices. We introduce AdaBet, a gradient-free layer selection approach to rank important layers, followed by important channels of these layers, by analyzing topological features of their activation spaces through Betti Numbers and using forward passes alone. AdaBet allows selecting layers and channels with high learning capacity, which are important for retraining and adaptation, without requiring labels or gradients. Evaluating AdaBet on sixteen pairs of benchmark models and datasets shows Ad-aBet achieves an average gain of 2.5% more classification accuracy over gradient-based baselines while reducing average peak memory consumption by 40%. We open-source our code at https://github.com/Nokia-Bell- Labs/efficient_layer_selection
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext efe0b3f7-1792-48ad-8967-795aa2040432Builds on10
- Robust Speech Recognition via Large-Scale Weak SupervisionAlec Radford, Jong Wook Kim, Tao Xu, Greg Brockman et al.ICML 2023 · 6,966 citations
- TinyTL: Reduce Memory, Not Parameters for Efficient On-Device LearningHan Cai, Chuang Gan, Ligeng Zhu, Song HanNeurIPS 2020 · 375 citations
- On-Device Training Under 256KB MemoryJi Lin, Ligeng Zhu, Wei-Ming Chen, Wei-Chen Wang et al.NeurIPS 2022 · 345 citations
- Surgical Fine-Tuning Improves Adaptation to Distribution ShiftsYoonho Lee, Annie S. Chen, Fahim Tajwar, Ananya Kumar et al.ICLR 2023 · 47 citations
- Understanding Approximate Fisher Information for Fast Convergence of Natural Gradient Descent in Wide Neural NetworksRyo Karakida, Kazuki OsawaNeurIPS 2020 · 39 citations
Related papers
- Study of Training Dynamics for Memory-Constrained Fine-TuningAël Quélennec, Nour Hezbri, Pavlo Mozharovskyi, Van-Tam Nguyen et al.ICLR 2026 · 1 citation
- SURGEON: Memory-Adaptive Fully Test-Time Adaptation via Dynamic Activation SparsityKe Ma, Jiaqi Tang, Bin Guo, Fan Dang et al.CVPR 2025
- RepNet: Efficient On-Device Learning via Feature ReprogrammingLi Yang, Adnan Siraj Rakin, Deliang FanCVPR 2022 · 18 citations
- EdgeFormer: A Parameter-Efficient Transformer for On-Device Seq2seq GenerationTao Ge, Si-Qing Chen, Furu WeiEMNLP 2022 · 16 citations
- Enabling On-Tiny-Device Model Personalization via Gradient Condensing and Alternant Partial UpdateZhenge Jia, Yiyang Shi, Zeyu Bao, Zirui Wang et al.DAC 2025
