SC2021Top-tier venue
Clairvoyant prefetching for distributed machine learning I/O
Nikoli Dryden, Roman Böhringer, Tal Ben-Nun, Torsten Hoefler
Abstract
I/O is emerging as a major bottleneck for machine learning training, especially in distributed environments. Indeed, at large scale, I/O takes as much as 85% of training time. Addressing this I/O bottleneck necessitates careful optimization, as optimal data ingestion pipelines differ between systems, and require a delicate balance between access to local storage, external filesystems, and remote nodes. We introduce NoPFS, a machine learning I/O middleware, which provides a scalable, flexible, and easy-to-use solution to the I/O bottleneck. NoPFS uses clairvoyance: Given the seed generating the random access pattern for training with SGD, it can exactly predict when and where a sample will be accessed. We combine this with an analysis of access patterns and a performance model to provide distributed caching policies that adapt to different datasets and storage hierarchies. NoPFS reduces I/O times and improves end-to-end training by up to 5.4× on the ImageNet-1k, ImageNet-22k, and CosmoFlow datasets.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers10
- SHADE: Enable Fundamental Cacheability for Distributed Deep Learning TrainingRedwan Ibne Seraj Khan, Ahmad Hossein Yazdani, Yuqi Fu, Arnab K. Paul et al.FAST 2023 · 29 citations
- Pecan: Cost-Efficient ML Data Preprocessing with Automatic Transformation Ordering and Hybrid PlacementDan Graur, Oto Mraz, Muyu Li, Mohammad Sepehr Pourghannad et al.USENIX ATC 2024 · 18 citations
- GPU-Enabled Asynchronous Multi-level Checkpoint Caching and PrefetchingAvinash Maurya, M. Mustafa Rafique, Thierry Tonellot, Hussain J. AlSalem et al.HPDC 2023 · 13 citations
- KAKURENBO: Adaptively Hiding Samples in Deep Neural Network TrainingTruong Thao Nguyen, Balazs Gerofi, Edgar Josafat Martinez-Noriega, François Trahay et al.NeurIPS 2023 · 6 citations
- Compressing multidimensional weather and climate data into neural networksLangwen Huang, Torsten HoeflerICLR 2023 · 6 citations
Related papers
- FalconFS: Distributed File System for Large-Scale Deep Learning PipelineJingwei Xu, Junbin Kang, Mingkai Dong, Mingyu Liu et al.NSDI 2026 · 2 citations
- FastFlow: Accelerating Deep Learning Model Training with Smart Offloading of Input Data PipelineTaegeon Um, Byungsoo Oh, Byeongchan Seo, Minhyeok Kweun et al.VLDB 2023 · 45 citations
- NDPipe: Exploiting Near-data Processing for Scalable Inference and Continuous Training in Photo StorageJungwoo Kim, Seonggyun Oh, Jaeha Kung, Yeseong Kim et al.ASPLOS 2024
- FFCV: Accelerating Training by Removing Data BottlenecksGuillaume Leclerc, Andrew Ilyas, Logan Engstrom, Sung Min Park et al.CVPR 2023
- Cachew: Machine Learning Input Data Processing as a ServiceDan Graur, Damien Aymon, Dan Kluser, Tanguy Albrici et al.USENIX ATC 2022 · 43 citations
