Egeria: Efficient DNN Training with Knowledge-Guided Layer Freezing
Yiding Wang, Decang Sun, Kai Chen, Fan Lai, Mosharaf Chowdhury
Abstract
Training deep neural networks (DNNs) is time-consuming. While most existing solutions try to overlap/schedule computation and communication for efficient training, this paper goes one step further by skipping computing and communication through DNN layer freezing. Our key insight is that the training progress of internal DNN layers differs significantly, and front layers often become well-trained much earlier than deep layers. To explore this, we first introduce the notion of training plasticity to quantify the training progress of internal DNN layers. Then we design Egeria, a knowledgeguided DNN training system that employs semantic knowledge from a reference model to accurately evaluate individual layers' training plasticity and safely freeze the converged ones, saving their corresponding backward computation and communication. Our reference model is generated on the fly using quantization techniques and runs forward operations asynchronously on available CPUs to minimize the overhead. In addition, Egeria caches the intermediate outputs of the frozen layers with prefetching to further skip the forward computation. Our implementation and testbed experiments with popular vision and language models show that Egeria achieves 19%-43% training speedup w.r.t. the state-of-the-art without sacrificing accuracy. 1 Plasticity quantifies a layer's training progress toward convergence, which is borrowed from neuroplasticity in neural science and child development [15]. Basically, a DNN layer's training plasticity will gradually decrease and become stable as it converges. 1
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 1d338c1a-ef71-4b08-92c3-99ff15a9a88aCited by top-tier papers9
- CacheGen: KV Cache Compression and Streaming for Fast Large Language Model ServingYuhan Liu, Hanchen Li, Yihua Cheng, Siddhant Ray et al.SIGCOMM 2024 · 111 citations
- AdaEmbed: Adaptive Embedding for Large-Scale Recommendation ModelsFan Lai, Wei Zhang, Rui Liu, William Tsai et al.OSDI 2023 · 23 citations
- Balanced and Elastic End-to-end Training of Dynamic LLMsMohamed Wahib, Muhammed Abdullah Soyturk, Didem UnatSC 2025 · 3 citations
- NeuroFlux: Memory-Efficient CNN Training Using Adaptive Local LearningDhananjay Saikumar, Blesson VargheseEuroSys 2024 · 2 citations
- Breaking the Memory Wall for Heterogeneous Federated Learning via Progressive TrainingYebo Wu, Li Li, Cheng-Zhong XuKDD 2025 · 2 citations
Builds on16
- Similarity-Preserving Knowledge DistillationFrederick Tung, Greg MoriICCV 2019 · 1,214 citations
- MCUNet: Tiny Deep Learning on IoT DevicesJi Lin, Wei-Ming Chen, Yujun Lin, John Cohn et al.NeurIPS 2020 · 827 citations
- A Unified Architecture for Accelerating Distributed DNN Training in Heterogeneous GPU/CPU ClustersYimin Jiang, Yibo Zhu, Chang Lan, Bairen Yi et al.OSDI 2020 · 390 citations
- FedScale: Benchmarking Model and System Performance of Federated Learning at ScaleFan Lai, Yinwei Dai, Sanjay Sri Vallabh Singapuram, Jiachen Liu et al.ICML 2022 · 280 citations
- Knowledge Distillation from Internal RepresentationsGustavo Aguilar, Yuan Ling, Yu Zhang, Benjamin Z. Yao et al.AAAI 2020 · 199 citations
Related papers
- Convergence-Aware Neural Network TrainingHyungjun Oh, Yongseung Yu, Giha Ryu, Gunjoo Ahn et al.DAC 2020 · 4 citations
- SmartFRZ: An Efficient Training Framework using Attention-Based Layer FreezingSheng Li, Geng Yuan, Yue Dai, Youtao Zhang et al.ICLR 2023 · 1 citation
- ADA-GP: Accelerating DNN Training By Adaptive Gradient PredictionVahid Janfaza, Shantanu Mandal, Farabi Mahmud, Abdullah MuzahidMICRO 2023 · 3 citations
- Training Acceleration for Deep Neural Networks: A Hybrid Parallelization StrategyZihao Zeng, Chubo Liu, Zhuo Tang, Wanli Chang et al.DAC 2021 · 13 citations
- Demand Layering for Real-Time DNN Inference with Minimized Memory UsageMingoo Ji, Saehanseul Yi, Changjin Koo, Sol Ahn et al.RTSS 2022 · 21 citations
