Quiver: An Informed Storage Cache for Deep Learning
Abhishek Vijaya Kumar, Muthian Sivathanu
摘要
We introduce Quiver, an informed storage cache for deep learning training (DLT) jobs in a cluster of GPUs. Quiver employs domain-specific intelligence within the caching layer, to achieve much higher efficiency compared to a generic storage cache. First, Quiver uses a secure hash-based addressing to transparently reuse cached data across multiple jobs and even multiple users operating on the same dataset. Second, by co-designing with the deep learning framework (e.g. PyTorch), Quiver employs a technique of substitutable cache hits to get more value from the existing contents of the cache, thus avoiding cache thrashing when cache capacity is much smaller than the working set. Third, Quiver dynamically prioritizes cache allocation to jobs that benefit the most from the caching. With a prototype implementation in PyTorch, we show that Quiver can significantly improve throughput of deep learning workloads.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper24
- Analyzing and Mitigating Data Stalls in DNN TrainingJayashree Mohan, Amar Phanishayee, Ashish Raniwala, Vijay ChidambaramVLDB 2021 · 被引用 142 次
- Accelerating Recommendation System Training by Leveraging Popular ChoicesMuhammad Adnan, Yassaman Ebrahimzadeh Maboud, Divya Mahajan, Prashant J. NairVLDB 2022 · 被引用 70 次
- FlashNeuron: SSD-Enabled Large-Batch Training of Very Deep Neural NetworksJonghyun Bae, Jongsung Lee, Yunho Jin, Sam Son 等FAST 2021 · 被引用 64 次
- FastFlow: Accelerating Deep Learning Model Training with Smart Offloading of Input Data PipelineTaegeon Um, Byungsoo Oh, Byeongchan Seo, Minhyeok Kweun 等VLDB 2023 · 被引用 45 次
- Cachew: Machine Learning Input Data Processing as a ServiceDan Graur, Damien Aymon, Dan Kluser, Tanguy Albrici 等USENIX ATC 2022 · 被引用 43 次
相关 Paper
- SHADE: Enable Fundamental Cacheability for Distributed Deep Learning TrainingRedwan Ibne Seraj Khan, Ahmad Hossein Yazdani, Yuqi Fu, Arnab K. Paul 等FAST 2023 · 被引用 29 次
- iCache: An Importance-Sampling-Informed Cache for Accelerating I/O-Bound DNN Model TrainingWeijian Chen, Shuibing He, Yaowen Xu, Xuechen Zhang 等HPCA 2023 · 被引用 20 次
- SiloD: A Co-design of Caching and Scheduling for Deep Learning ClustersHanyu Zhao, Zhenhua Han, Zhi Yang, Quanlu Zhang 等EuroSys 2023 · 被引用 22 次
- Dynamic Resource Allocation for Deep Learning Clusters with Separated Compute and StorageMingxia Li, Zhenhua Han, Chi Zhang, Ruiting Zhou 等INFOCOM 2023 · 被引用 3 次
- Multi-resource interleaving for deep learning trainingYihao Zhao, Yuanqiang Liu, Yanghua Peng, Yibo Zhu 等SIGCOMM 2022 · 被引用 78 次
