iCache: An Importance-Sampling-Informed Cache for Accelerating I/O-Bound DNN Model Training
Weijian Chen, Shuibing He, Yaowen Xu, Xuechen Zhang, Siling Yang, Shuang Hu, Xian-He Sun, Gang Chen
摘要
Fetching a large amount of DNN training data from storage systems incurs long I/O latency and fetch stalls of GPUs. Importance sampling in DNN training can reduce the amount of data computing on GPUs while maintaining a similar model accuracy. However, existing DNN training frameworks do not have a cache layer that reduces the number of data fetches and manages cached items according to sample importance, resulting in unnecessary data fetches, poor cache hit ratios, and random I/Os when importance sampling is used.
In this paper, we design a new importance-sampling-informed cache, namely, iCACHE, to accelerate I/O bound DNN training jobs. iCACHE only fetches parts of samples instead of all samples in the dataset. The cache is partitioned into two regions: Hcache and L-cache, which store samples of high importance and low importance respectively. Rather than using recency or frequency, we manage data items in H-cache according to their corresponding sample importance. When there is a cache miss in L-cache, we use sample substitutability and dynamic packaging to improve the cache hit ratio and reduce the number of random I/Os. When multiple concurrent jobs access the same datasets in H-cache, we design a model to assign the relative importance values to cached samples to avoid cache thrashing, which may happen when there is no coordination among the concurrent training jobs. Our experimental results show that iCACHE has a negligible impact on training accuracy and speeds up the DNN training time by up to 2.0⇥ compared to the state-of-the-art caching systems.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- SHADE: Enable Fundamental Cacheability for Distributed Deep Learning TrainingRedwan Ibne Seraj Khan, Ahmad Hossein Yazdani, Yuqi Fu, Arnab K. Paul 等FAST 2023 · 被引用 29 次
- Pecan: Cost-Efficient ML Data Preprocessing with Automatic Transformation Ordering and Hybrid PlacementDan Graur, Oto Mraz, Muyu Li, Mohammad Sepehr Pourghannad 等USENIX ATC 2024 · 被引用 18 次
- cedar: Optimized and Unified Machine Learning Input Data PipelinesMark Zhao, Emanuel Adamiak, Christos KozyrakisVLDB 2025 · 被引用 13 次
- PreSto: An In-Storage Data Preprocessing System for Training Recommendation ModelsYunjae Lee, Hyeseong Kim, Minsoo RhuISCA 2024 · 被引用 8 次
它引用的顶会 Paper5
- Analyzing and Mitigating Data Stalls in DNN TrainingJayashree Mohan, Amar Phanishayee, Ashish Raniwala, Vijay ChidambaramVLDB 2021 · 被引用 142 次
- Quiver: An Informed Storage Cache for Deep LearningAbhishek Vijaya Kumar, Muthian SivathanuFAST 2020 · 被引用 91 次
- NVAlloc: rethinking heap metadata management in persistent memory allocatorsZheng Dang, Shuibing He, Peiyi Hong, Zhenxin Li 等ASPLOS 2022 · 被引用 35 次
- Refurbish Your Training Data: Reusing Partially Augmented Samples for Faster Deep Neural Network TrainingGyewon Lee, Irene Lee, Hyeonmin Ha, Kyung-Geun Lee 等USENIX ATC 2021 · 被引用 25 次
- XPGraph: XPline-Friendly Persistent Memory Graph Stores for Large-Scale Evolving GraphsRui Wang, Shuibing He, Weixu Zong, Yongkun Li 等MICRO 2022 · 被引用 21 次
相关 Paper
- A Deep Learning Dataloader with Shared Data PreparationJian Xie, Jingwei Xu, Guochang Wang, Yuan Yao 等NeurIPS 2022 · 被引用 8 次
- Preparation Meets Opportunity: Enhancing Data Preprocessing for ML Training With SenecaOmkar Desai, Ziyang Jiao, Shuyi Pei, Janki Bhimani 等FAST 2026 · 被引用 3 次
- HyCache: Hybrid Caching for Accelerating DNN Input Preprocessing PipelinesKeshav Vinayak Jha, Shweta Pandey, Murali Annavaram, Arkaprava BasuUSENIX ATC 2025 · 被引用 2 次
- DiskGNN: Bridging I/O Efficiency and Model Accuracy for Out-of-Core GNN TrainingRenjie Liu, Yichuan Wang, Xiao Yan, Haitian Jiang 等SIGMOD 2025 · 被引用 8 次
- Zico: Efficient GPU Memory Sharing for Concurrent DNN TrainingGangmuk Lim, Jeongseob Ahn, Wencong Xiao, Youngjin Kwon 等USENIX ATC 2021 · 被引用 65 次
