Study of Training Dynamics for Memory-Constrained Fine-Tuning
Aël Quélennec, Nour Hezbri, Pavlo Mozharovskyi, Van-Tam Nguyen, Enzo Tartaglione
Abstract
Memory-efficient training of deep neural networks has become increasingly important as models grow larger while deployment environments impose strict resource constraints. We propose TraDy, a novel transfer learning scheme leveraging two key insights: layer importance for updates is architecture-dependent and determinable a priori, while dynamic stochastic channel selection provides superior gradient approximation compared to static approaches. We introduce a dynamic channel selection approach that stochastically resamples channels between epochs within preselected layers. Extensive experiments demonstrate TraDy achieves state-of-the-art performance across various downstream tasks and architectures while maintaining strict memory constraints, achieving up to 99% activation sparsity, 95% weight derivative sparsity, and 97% reduction in FLOPs for weight derivative computation.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext dc4aea5b-d4ff-4b4c-9807-051f1cb3e6f9Builds on17
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- MCUNet: Tiny Deep Learning on IoT DevicesJi Lin, Wei-Ming Chen, Yujun Lin, John Cohn et al.NeurIPS 2020 · 827 citations
- TinyTL: Reduce Memory, Not Parameters for Efficient On-Device LearningHan Cai, Chuang Gan, Ligeng Zhu, Song HanNeurIPS 2020 · 375 citations
- On-Device Training Under 256KB MemoryJi Lin, Ligeng Zhu, Wei-Ming Chen, Wei-Chen Wang et al.NeurIPS 2022 · 345 citations
Related papers
- SURGEON: Memory-Adaptive Fully Test-Time Adaptation via Dynamic Activation SparsityKe Ma, Jiaqi Tang, Bin Guo, Fan Dang et al.CVPR 2025
- Training Neural Networks with Fixed Sparse MasksYi-Lin Sung, Varun Nair, Colin RaffelNeurIPS 2021 · 295 citations
- AdaBet: Gradient-free Layer Selection for Efficient Training of Deep Neural NetworksIrene Tenison, Soumyajit Chatterjee, Fahim Kawsar, Mohammad MalekzadehCVPR 2026
- Finding trainable sparse networks through Neural Tangent TransferTianlin Liu, Friedemann ZenkeICML 2020 · 40 citations
- SparseProp: Efficient Sparse Backpropagation for Faster Training of Neural Networks at the EdgeMahdi Nikdan, Tommaso Pegolotti, Eugenia Iofinova, Eldar Kurtic et al.ICML 2023 · 14 citations
