SiloD: A Co-design of Caching and Scheduling for Deep Learning Clusters
Hanyu Zhao, Zhenhua Han, Zhi Yang, Quanlu Zhang, Mingxia Li, Fan Yang, Qianxi Zhang, Binyang Li, Yuqing Yang, Lili Qiu, Lintao Zhang, Lidong Zhou
Abstract
Deep learning training on cloud platforms usually follows the tradition of the separation of storage and computing. The training executes on a compute cluster equipped with GPUs/TPUs while reading data from a separate cluster hosting the storage service. To alleviate the potential bottleneck, a training cluster usually leverages its local storage as a cache to reduce the remote IO from the storage cluster. However, existing deep learning schedulers do not manage storage resources thus fail to consider the diverse caching effects across different training jobs. This could degrade scheduling quality significantly.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 58634757-5e66-455c-9168-867f00ea21c2Cited by top-tier papers4
- An Empirical Study on Low GPU Utilization of Deep Learning JobsYanjie Gao, Yichen He, Xinze Li, Bo Zhao et al.ICSE 2024 · 22 citations
- cedar: Optimized and Unified Machine Learning Input Data PipelinesMark Zhao, Emanuel Adamiak, Christos KozyrakisVLDB 2025 · 13 citations
- GPEmu: A GPU Emulator for Faster and Cheaper Prototyping and Evaluation of Deep Learning System ResearchMeng Wang, Gus Waldspurger, Naufal Ananda, Yuyang Huang et al.VLDB 2025 · 1 citation
- Arena: Efficiently Training Large Models via Dynamic Scheduling and Adaptive Parallelism Co-DesignChunyu Xue, Weihao Cui, Quan Chen, Chen Chen et al.EuroSys 2026
Related papers
- Dynamic Resource Allocation for Deep Learning Clusters with Separated Compute and StorageMingxia Li, Zhenhua Han, Chi Zhang, Ruiting Zhou et al.INFOCOM 2023 · 3 citations
- Multi-resource interleaving for deep learning trainingYihao Zhao, Yuanqiang Liu, Yanghua Peng, Yibo Zhu et al.SIGCOMM 2022 · 78 citations
- OS2G: A High-Performance DPU Offloading Architecture for GPU-based Deep Learning with Object StorageZhen Jin, Yiquan Chen, Mingxu Liang, Yijing Wang et al.ASPLOS 2025 · 5 citations
- Quiver: An Informed Storage Cache for Deep LearningAbhishek Vijaya Kumar, Muthian SivathanuFAST 2020 · 91 citations
- Heet: Accelerating Elastic Training in Heterogeneous Deep Learning ClustersZizhao Mo, Huanle Xu, Chengzhong XuASPLOS 2024 · 22 citations
