SHADE: Enable Fundamental Cacheability for Distributed Deep Learning Training
Redwan Ibne Seraj Khan, Ahmad Hossein Yazdani, Yuqi Fu, Arnab K. Paul, Bo Ji, Xun Jian, Yue Cheng, Ali Raza Butt
摘要
Deep learning training (DLT) applications exhibit unique I/O workload behaviors that pose new challenges for storage system design. DLT is I/O intensive since data samples need to be fetched continuously from a remote storage. Accelerators such as GPUs have been extensively used to support these applications. As accelerators become more powerful and more data-hungry, the I/O performance lags behind. This creates a crucial performance bottleneck, especially in distributed DLT. At the same time, the exponentially growing dataset sizes make it impossible to store these datasets entirely in memory. While today's DLT frameworks typically use a random sampling policy that treat all samples uniformly equally, recent findings indicate that not all samples are equally important and different data samples contribute differently towards improving the accuracy of a model. This observation creates an opportunity for DLT I/O optimizations by exploiting the data locality enabled by importance sampling.
To this end, we design and implement SHADE, a new DLTaware caching system that detects fine-grained importance variations at per-sample level and leverages the variance to make informed caching decisions for a distributed DLT job. SHADE adopts a novel, rank-based approach, which captures the relative importance of data samples across different minibatches. SHADE then dynamically updates the importance scores of all samples during training. With these techniques, SHADE manages to significantly improve the cache hit ratio of the DLT job, and thus, improves the job's training performance. Evaluation with representative computer vision (CV) models shows that SHADE, with a small cache, improves the cache hit ratio by up to 4.5× compared to the LRU caching policy.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper7
- Pecan: Cost-Efficient ML Data Preprocessing with Automatic Transformation Ordering and Hybrid PlacementDan Graur, Oto Mraz, Muyu Li, Mohammad Sepehr Pourghannad 等USENIX ATC 2024 · 被引用 18 次
- cedar: Optimized and Unified Machine Learning Input Data PipelinesMark Zhao, Emanuel Adamiak, Christos KozyrakisVLDB 2025 · 被引用 13 次
- Heimdall: Optimizing Storage I/O Admission with Extensive Machine Learning PipelineDaniar Heri Kurniawan, Rani Ayu Putri, Peiran Qin, Kahfi S. Zulkifli 等EuroSys 2025 · 被引用 3 次
- Preparation Meets Opportunity: Enhancing Data Preprocessing for ML Training With SenecaOmkar Desai, Ziyang Jiao, Shuyi Pei, Janki Bhimani 等FAST 2026 · 被引用 3 次
- GPEmu: A GPU Emulator for Faster and Cheaper Prototyping and Evaluation of Deep Learning System ResearchMeng Wang, Gus Waldspurger, Naufal Ananda, Yuyang Huang 等VLDB 2025 · 被引用 1 次
它引用的顶会 Paper7
- Lessons Learned from the Chameleon TestbedKate Keahey, Jason Anderson, Zhuo Zhen, Pierre Riteau 等USENIX ATC 2020 · 被引用 398 次
- Analyzing and Mitigating Data Stalls in DNN TrainingJayashree Mohan, Amar Phanishayee, Ashish Raniwala, Vijay ChidambaramVLDB 2021 · 被引用 142 次
- Quiver: An Informed Storage Cache for Deep LearningAbhishek Vijaya Kumar, Muthian SivathanuFAST 2020 · 被引用 91 次
- Preemptive All-reduce Scheduling for Expediting Distributed DNN TrainingYixin Bao, Yanghua Peng, Yangrui Chen, Chuan WuINFOCOM 2020 · 被引用 67 次
- Communication-Efficient Distributed Deep Learning with Merged Gradient Sparsification on GPUsShaohuai Shi, Qiang Wang, Xiaowen Chu, Bo Li 等INFOCOM 2020 · 被引用 66 次
相关 Paper
- iCache: An Importance-Sampling-Informed Cache for Accelerating I/O-Bound DNN Model TrainingWeijian Chen, Shuibing He, Yaowen Xu, Xuechen Zhang 等HPCA 2023 · 被引用 20 次
- SiloD: A Co-design of Caching and Scheduling for Deep Learning ClustersHanyu Zhao, Zhenhua Han, Zhi Yang, Quanlu Zhang 等EuroSys 2023 · 被引用 22 次
- Dynamic Resource Allocation for Deep Learning Clusters with Separated Compute and StorageMingxia Li, Zhenhua Han, Chi Zhang, Ruiting Zhou 等INFOCOM 2023 · 被引用 3 次
- A Deep Learning Dataloader with Shared Data PreparationJian Xie, Jingwei Xu, Guochang Wang, Yuan Yao 等NeurIPS 2022 · 被引用 8 次
- Elastic Resource Sharing for Distributed Deep LearningChangho Hwang, Taehyun Kim, Sunghyun Kim, Jinwoo Shin 等NSDI 2021 · 被引用 111 次
