HyCache: Hybrid Caching for Accelerating DNN Input Preprocessing Pipelines
Keshav Vinayak Jha, Shweta Pandey, Murali Annavaram, Arkaprava Basu
Abstract
End-to-end deep neural networks' (DNNs) training performance depends not only on the time spent in training the model weights but also on the time spent in loading and preprocessing the training data. Recent advances in GPU hardware have made training substantially faster. As a result, the bottleneck has shifted to the CPU-based input pipeline. This pipeline must fetch and transform each sample through multiple stages before it can be consumed by the GPU.
Prior works accelerate preprocessing by caching intermediate results across epochs, but suffer from several key limitations: 1 They cache either in memory or in storage, but are unable to leverage both together. 2 They can cache the output of a stage only if it can entirely fit in the cache, which is a severe limitation for larger datasets. 3 They can cache the output of only one of the stages which could be suboptimal.
We thus introduce Hybrid Cache (HyCache), a runtime that enables the caching of subsets of preprocessed data from multiple intermediate steps on both memory and storage. Hy-Cache possesses the ability to partially cache the outputs of a stage across both memory and storage. HyCache deploys integer linear programming (ILP) to automatically determine the best caching strategies across the memory and the storage by finding an optimal trade-off between recomputation and caching. Importantly, it does so without any manual intervention. HyCache outperforms state-of-the-art prior approaches, delivering a raw pipeline throughput improvement ranging in speedups from 1.11× to 10.1× over a variety of pipelines.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 682fc620-fab0-4cf2-883b-22c9ed5185f1Builds on8
- MobileViT: Light-weight, General-purpose, and Mobile-friendly Vision TransformerSachin Mehta, Mohammad RastegariICLR 2022 · 2,162 citations
- Analyzing and Mitigating Data Stalls in DNN TrainingJayashree Mohan, Amar Phanishayee, Ashish Raniwala, Vijay ChidambaramVLDB 2021 · 142 citations
- Quiver: An Informed Storage Cache for Deep LearningAbhishek Vijaya Kumar, Muthian SivathanuFAST 2020 · 91 citations
- Jointly Optimizing Preprocessing and Inference for DNN-based Visual AnalyticsDaniel Kang, Ankit Mathur, Teja Veeramacheneni, Peter Bailis et al.VLDB 2021 · 50 citations
- Cachew: Machine Learning Input Data Processing as a ServiceDan Graur, Damien Aymon, Dan Kluser, Tanguy Albrici et al.USENIX ATC 2022 · 43 citations
Related papers
- iCache: An Importance-Sampling-Informed Cache for Accelerating I/O-Bound DNN Model TrainingWeijian Chen, Shuibing He, Yaowen Xu, Xuechen Zhang et al.HPCA 2023 · 20 citations
- Where Is My Training Bottleneck? Hidden Trade-Offs in Deep Learning Preprocessing PipelinesAlexander Isenko, Ruben Mayer, Jeffrey Jedele, Hans-Arno JacobsenSIGMOD 2022 · 30 citations
- SiloD: A Co-design of Caching and Scheduling for Deep Learning ClustersHanyu Zhao, Zhenhua Han, Zhi Yang, Quanlu Zhang et al.EuroSys 2023 · 22 citations
- NDPipe: Exploiting Near-data Processing for Scalable Inference and Continuous Training in Photo StorageJungwoo Kim, Seonggyun Oh, Jaeha Kung, Yeseong Kim et al.ASPLOS 2024
- FastFlow: Accelerating Deep Learning Model Training with Smart Offloading of Input Data PipelineTaegeon Um, Byungsoo Oh, Byeongchan Seo, Minhyeok Kweun et al.VLDB 2023 · 45 citations
