Egeria: Efficient DNN Training with Knowledge-Guided Layer Freezing
Yiding Wang, Decang Sun, Kai Chen, Fan Lai, Mosharaf Chowdhury
摘要
Training deep neural networks (DNNs) is time-consuming. While most existing solutions try to overlap/schedule computation and communication for efficient training, this paper goes one step further by skipping computing and communication through DNN layer freezing. Our key insight is that the training progress of internal DNN layers differs significantly, and front layers often become well-trained much earlier than deep layers. To explore this, we first introduce the notion of training plasticity to quantify the training progress of internal DNN layers. Then we design Egeria, a knowledgeguided DNN training system that employs semantic knowledge from a reference model to accurately evaluate individual layers' training plasticity and safely freeze the converged ones, saving their corresponding backward computation and communication. Our reference model is generated on the fly using quantization techniques and runs forward operations asynchronously on available CPUs to minimize the overhead. In addition, Egeria caches the intermediate outputs of the frozen layers with prefetching to further skip the forward computation. Our implementation and testbed experiments with popular vision and language models show that Egeria achieves 19%-43% training speedup w.r.t. the state-of-the-art without sacrificing accuracy. 1 Plasticity quantifies a layer's training progress toward convergence, which is borrowed from neuroplasticity in neural science and child development [15]. Basically, a DNN layer's training plasticity will gradually decrease and become stable as it converges. 1
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper9
- CacheGen: KV Cache Compression and Streaming for Fast Large Language Model ServingYuhan Liu, Hanchen Li, Yihua Cheng, Siddhant Ray 等SIGCOMM 2024 · 被引用 111 次
- AdaEmbed: Adaptive Embedding for Large-Scale Recommendation ModelsFan Lai, Wei Zhang, Rui Liu, William Tsai 等OSDI 2023 · 被引用 23 次
- Balanced and Elastic End-to-end Training of Dynamic LLMsMohamed Wahib, Muhammed Abdullah Soyturk, Didem UnatSC 2025 · 被引用 3 次
- NeuroFlux: Memory-Efficient CNN Training Using Adaptive Local LearningDhananjay Saikumar, Blesson VargheseEuroSys 2024 · 被引用 2 次
- Breaking the Memory Wall for Heterogeneous Federated Learning via Progressive TrainingYebo Wu, Li Li, Cheng-Zhong XuKDD 2025 · 被引用 2 次
它引用的顶会 Paper16
- Similarity-Preserving Knowledge DistillationFrederick Tung, Greg MoriICCV 2019 · 被引用 1,214 次
- MCUNet: Tiny Deep Learning on IoT DevicesJi Lin, Wei-Ming Chen, Yujun Lin, John Cohn 等NeurIPS 2020 · 被引用 827 次
- A Unified Architecture for Accelerating Distributed DNN Training in Heterogeneous GPU/CPU ClustersYimin Jiang, Yibo Zhu, Chang Lan, Bairen Yi 等OSDI 2020 · 被引用 390 次
- FedScale: Benchmarking Model and System Performance of Federated Learning at ScaleFan Lai, Yinwei Dai, Sanjay Sri Vallabh Singapuram, Jiachen Liu 等ICML 2022 · 被引用 280 次
- Knowledge Distillation from Internal RepresentationsGustavo Aguilar, Yuan Ling, Yu Zhang, Benjamin Z. Yao 等AAAI 2020 · 被引用 199 次
相关 Paper
- Convergence-Aware Neural Network TrainingHyungjun Oh, Yongseung Yu, Giha Ryu, Gunjoo Ahn 等DAC 2020 · 被引用 4 次
- SmartFRZ: An Efficient Training Framework using Attention-Based Layer FreezingSheng Li, Geng Yuan, Yue Dai, Youtao Zhang 等ICLR 2023 · 被引用 1 次
- ADA-GP: Accelerating DNN Training By Adaptive Gradient PredictionVahid Janfaza, Shantanu Mandal, Farabi Mahmud, Abdullah MuzahidMICRO 2023 · 被引用 3 次
- Training Acceleration for Deep Neural Networks: A Hybrid Parallelization StrategyZihao Zeng, Chubo Liu, Zhuo Tang, Wanli Chang 等DAC 2021 · 被引用 13 次
- Demand Layering for Real-Time DNN Inference with Minimized Memory UsageMingoo Ji, Saehanseul Yi, Changjin Koo, Sol Ahn 等RTSS 2022 · 被引用 21 次
