USENIX ATC2022顶会
Cachew: Machine Learning Input Data Processing as a Service
Dan Graur, Damien Aymon, Dan Kluser, Tanguy Albrici, Chandramohan A. Thekkath, Ana Klimovic
摘要
Processing input data plays a vital role in ML training, impacting accuracy, throughput, and cost. The input pipeline, which is responsible for feeding data-hungry GPUs/TPUs with training examples, is a common bottleneck. Alleviating data stalls is critical yet challenging for users. While today's frameworks provide mechanisms to maximize input pipeline throughput (e.g., distributing data processing on remote CPU workers and/or reusing cached data transformations), leveraging these mechanisms to jointly optimize training time and cost is non-trivial. Users face two key challenges. First, ML schedulers focus on GPU/TPU resources, leaving users on their own to optimize multi-dimensional resource allocations for data processing. Second, input pipelines often consume excessive compute power to repeatedly transform the same data. Deciding which source or transformed data to cache is non-trivial: large datasets are expensive to store, the compute time saved by caching is not always the bottleneck for end-toend training, and transformations may not be deterministic, hence reusing transformed data can impact accuracy.
We propose Cachew, a fully-managed service for ML data processing. Cachew dynamically scales distributed resources for data processing to avoid stalls in training jobs. The service also automatically applies caching when and where it is performance/cost-effective to reuse preprocessed data within and across jobs. Our key contributions are autoscaling and autocaching policies, which leverage domain-specific metrics collected at data workers and training clients (rather than generic resource utilization metrics) to minimize training time and cost. Compared to scaling workers with Kubernetes, Cachew's policies reduce training time by up to 4.1× and training cost by 1.1× to 3.8×.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper17
- Orion: Interference-aware, Fine-grained GPU Sharing for ML ApplicationsFoteini Strati, Xianzhe Ma, Ana KlimovicEuroSys 2024 · 被引用 96 次
- FastFlow: Accelerating Deep Learning Model Training with Smart Offloading of Input Data PipelineTaegeon Um, Byungsoo Oh, Byeongchan Seo, Minhyeok Kweun 等VLDB 2023 · 被引用 45 次
- Pecan: Cost-Efficient ML Data Preprocessing with Automatic Transformation Ordering and Hybrid PlacementDan Graur, Oto Mraz, Muyu Li, Mohammad Sepehr Pourghannad 等USENIX ATC 2024 · 被引用 18 次
- cedar: Optimized and Unified Machine Learning Input Data PipelinesMark Zhao, Emanuel Adamiak, Christos KozyrakisVLDB 2025 · 被引用 13 次
- Tectonic-Shift: A Composite Storage Fabric for Large-Scale ML TrainingMark Zhao, Satadru Pan, Niket Agarwal, Zhaoduo Wen 等USENIX ATC 2023 · 被引用 13 次
它引用的顶会 Paper12
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 被引用 24,064 次
- CutMix: Regularization Strategy to Train Strong Classifiers With Localizable FeaturesSangdoo Yun, Dongyoon Han, Sanghyuk Chun, Seong Joon Oh 等ICCV 2019 · 被引用 5,843 次
- RandAugment: Practical Automated Data Augmentation with a Reduced Search SpaceEkin Dogus Cubuk, Barret Zoph, Jonathon Shlens, Quoc LeNeurIPS 2020 · 被引用 4,453 次
- EfficientNetV2: Smaller Models and Faster TrainingMingxing Tan, Quoc V. LeICML 2021 · 被引用 4,239 次
- Big Self-Supervised Models are Strong Semi-Supervised LearnersTing Chen, Simon Kornblith, Kevin Swersky, Mohammad Norouzi 等NeurIPS 2020 · 被引用 2,611 次
相关 Paper
- Analyzing and Mitigating Data Stalls in DNN TrainingJayashree Mohan, Amar Phanishayee, Ashish Raniwala, Vijay ChidambaramVLDB 2021 · 被引用 142 次
- Dynamic Resource Allocation for Deep Learning Clusters with Separated Compute and StorageMingxia Li, Zhenhua Han, Chi Zhang, Ruiting Zhou 等INFOCOM 2023 · 被引用 3 次
- SiloD: A Co-design of Caching and Scheduling for Deep Learning ClustersHanyu Zhao, Zhenhua Han, Zhi Yang, Quanlu Zhang 等EuroSys 2023 · 被引用 22 次
- HyCache: Hybrid Caching for Accelerating DNN Input Preprocessing PipelinesKeshav Vinayak Jha, Shweta Pandey, Murali Annavaram, Arkaprava BasuUSENIX ATC 2025 · 被引用 2 次
- Preparation Meets Opportunity: Enhancing Data Preprocessing for ML Training With SenecaOmkar Desai, Ziyang Jiao, Shuyi Pei, Janki Bhimani 等FAST 2026 · 被引用 3 次
