iCache: An Importance-Sampling-Informed Cache for Accelerating I/O-Bound DNN Model Training
Weijian Chen, Shuibing He, Yaowen Xu, Xuechen Zhang, Siling Yang, Shuang Hu, Xian-He Sun, Gang Chen
Abstract
Fetching a large amount of DNN training data from storage systems incurs long I/O latency and fetch stalls of GPUs. Importance sampling in DNN training can reduce the amount of data computing on GPUs while maintaining a similar model accuracy. However, existing DNN training frameworks do not have a cache layer that reduces the number of data fetches and manages cached items according to sample importance, resulting in unnecessary data fetches, poor cache hit ratios, and random I/Os when importance sampling is used.
In this paper, we design a new importance-sampling-informed cache, namely, iCACHE, to accelerate I/O bound DNN training jobs. iCACHE only fetches parts of samples instead of all samples in the dataset. The cache is partitioned into two regions: Hcache and L-cache, which store samples of high importance and low importance respectively. Rather than using recency or frequency, we manage data items in H-cache according to their corresponding sample importance. When there is a cache miss in L-cache, we use sample substitutability and dynamic packaging to improve the cache hit ratio and reduce the number of random I/Os. When multiple concurrent jobs access the same datasets in H-cache, we design a model to assign the relative importance values to cached samples to avoid cache thrashing, which may happen when there is no coordination among the concurrent training jobs. Our experimental results show that iCACHE has a negligible impact on training accuracy and speeds up the DNN training time by up to 2.0⇥ compared to the state-of-the-art caching systems.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 3dd2bb91-172c-4306-9346-19159455b45fCited by top-tier papers4
- SHADE: Enable Fundamental Cacheability for Distributed Deep Learning TrainingRedwan Ibne Seraj Khan, Ahmad Hossein Yazdani, Yuqi Fu, Arnab K. Paul et al.FAST 2023 · 29 citations
- Pecan: Cost-Efficient ML Data Preprocessing with Automatic Transformation Ordering and Hybrid PlacementDan Graur, Oto Mraz, Muyu Li, Mohammad Sepehr Pourghannad et al.USENIX ATC 2024 · 18 citations
- cedar: Optimized and Unified Machine Learning Input Data PipelinesMark Zhao, Emanuel Adamiak, Christos KozyrakisVLDB 2025 · 13 citations
- PreSto: An In-Storage Data Preprocessing System for Training Recommendation ModelsYunjae Lee, Hyeseong Kim, Minsoo RhuISCA 2024 · 8 citations
Builds on5
- Analyzing and Mitigating Data Stalls in DNN TrainingJayashree Mohan, Amar Phanishayee, Ashish Raniwala, Vijay ChidambaramVLDB 2021 · 142 citations
- Quiver: An Informed Storage Cache for Deep LearningAbhishek Vijaya Kumar, Muthian SivathanuFAST 2020 · 91 citations
- NVAlloc: rethinking heap metadata management in persistent memory allocatorsZheng Dang, Shuibing He, Peiyi Hong, Zhenxin Li et al.ASPLOS 2022 · 35 citations
- Refurbish Your Training Data: Reusing Partially Augmented Samples for Faster Deep Neural Network TrainingGyewon Lee, Irene Lee, Hyeonmin Ha, Kyung-Geun Lee et al.USENIX ATC 2021 · 25 citations
- XPGraph: XPline-Friendly Persistent Memory Graph Stores for Large-Scale Evolving GraphsRui Wang, Shuibing He, Weixu Zong, Yongkun Li et al.MICRO 2022 · 21 citations
Related papers
- A Deep Learning Dataloader with Shared Data PreparationJian Xie, Jingwei Xu, Guochang Wang, Yuan Yao et al.NeurIPS 2022 · 8 citations
- Preparation Meets Opportunity: Enhancing Data Preprocessing for ML Training With SenecaOmkar Desai, Ziyang Jiao, Shuyi Pei, Janki Bhimani et al.FAST 2026 · 3 citations
- HyCache: Hybrid Caching for Accelerating DNN Input Preprocessing PipelinesKeshav Vinayak Jha, Shweta Pandey, Murali Annavaram, Arkaprava BasuUSENIX ATC 2025 · 2 citations
- DiskGNN: Bridging I/O Efficiency and Model Accuracy for Out-of-Core GNN TrainingRenjie Liu, Yichuan Wang, Xiao Yan, Haitian Jiang et al.SIGMOD 2025 · 8 citations
- Zico: Efficient GPU Memory Sharing for Concurrent DNN TrainingGangmuk Lim, Jeongseob Ahn, Wencong Xiao, Youngjin Kwon et al.USENIX ATC 2021 · 65 citations
