Clairvoyant prefetching for distributed machine learning I/O
Nikoli Dryden, Roman Böhringer, Tal Ben-Nun, Torsten Hoefler
摘要
I/O is emerging as a major bottleneck for machine learning training, especially in distributed environments. Indeed, at large scale, I/O takes as much as 85% of training time. Addressing this I/O bottleneck necessitates careful optimization, as optimal data ingestion pipelines differ between systems, and require a delicate balance between access to local storage, external filesystems, and remote nodes. We introduce NoPFS, a machine learning I/O middleware, which provides a scalable, flexible, and easy-to-use solution to the I/O bottleneck. NoPFS uses clairvoyance: Given the seed generating the random access pattern for training with SGD, it can exactly predict when and where a sample will be accessed. We combine this with an analysis of access patterns and a performance model to provide distributed caching policies that adapt to different datasets and storage hierarchies. NoPFS reduces I/O times and improves end-to-end training by up to 5.4× on the ImageNet-1k, ImageNet-22k, and CosmoFlow datasets.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper10
- SHADE: Enable Fundamental Cacheability for Distributed Deep Learning TrainingRedwan Ibne Seraj Khan, Ahmad Hossein Yazdani, Yuqi Fu, Arnab K. Paul 等FAST 2023 · 被引用 29 次
- Pecan: Cost-Efficient ML Data Preprocessing with Automatic Transformation Ordering and Hybrid PlacementDan Graur, Oto Mraz, Muyu Li, Mohammad Sepehr Pourghannad 等USENIX ATC 2024 · 被引用 18 次
- GPU-Enabled Asynchronous Multi-level Checkpoint Caching and PrefetchingAvinash Maurya, M. Mustafa Rafique, Thierry Tonellot, Hussain J. AlSalem 等HPDC 2023 · 被引用 13 次
- KAKURENBO: Adaptively Hiding Samples in Deep Neural Network TrainingTruong Thao Nguyen, Balazs Gerofi, Edgar Josafat Martinez-Noriega, François Trahay 等NeurIPS 2023 · 被引用 6 次
- Compressing multidimensional weather and climate data into neural networksLangwen Huang, Torsten HoeflerICLR 2023 · 被引用 6 次
相关 Paper
- FalconFS: Distributed File System for Large-Scale Deep Learning PipelineJingwei Xu, Junbin Kang, Mingkai Dong, Mingyu Liu 等NSDI 2026 · 被引用 2 次
- FastFlow: Accelerating Deep Learning Model Training with Smart Offloading of Input Data PipelineTaegeon Um, Byungsoo Oh, Byeongchan Seo, Minhyeok Kweun 等VLDB 2023 · 被引用 45 次
- NDPipe: Exploiting Near-data Processing for Scalable Inference and Continuous Training in Photo StorageJungwoo Kim, Seonggyun Oh, Jaeha Kung, Yeseong Kim 等ASPLOS 2024
- FFCV: Accelerating Training by Removing Data BottlenecksGuillaume Leclerc, Andrew Ilyas, Logan Engstrom, Sung Min Park 等CVPR 2023
- Cachew: Machine Learning Input Data Processing as a ServiceDan Graur, Damien Aymon, Dan Kluser, Tanguy Albrici 等USENIX ATC 2022 · 被引用 43 次
