Pecan: Cost-Efficient ML Data Preprocessing with Automatic Transformation Ordering and Hybrid Placement
Dan Graur, Oto Mraz, Muyu Li, Mohammad Sepehr Pourghannad, Chandramohan A. Thekkath, Ana Klimovic
Abstract
Input data preprocessing is a common bottleneck in machine learning (ML) jobs, that can significantly increase training time and cost as expensive GPUs or TPUs idle waiting for input data. Previous work has shown that offloading data preprocessing to remote CPU servers successfully alleviates data stalls and improves training time. However, remote CPU workers in disaggregated data processing systems comprise a significant fraction of total training costs. Meanwhile, current disaggregated solutions often underutilize CPU and DRAM resources available on ML accelerator nodes. We propose two approaches to alleviate ML input data stalls while minimizing costs. First, we dynamically schedule data preprocessing workers on ML accelerator host resources to minimize the number of remote CPU workers needed to achieve peak data ingestion bandwidth. Second, we analyze the characteristics of input pipelines and automatically reorder transformations to increase data preprocessing worker throughput. We observe that relaxing commutativity increases throughput while maintaining high model accuracy for a variety of ML data pipelines. We build Pecan, an ML data preprocessing service that automates data preprocessing worker placement and transformation reordering decisions. Pecan reduces preprocessing costs by 87% on average and total training costs by up to 60% compared to training with state-of-the-art disaggregated data preprocessing and total training costs by 55% on average compared to collocated data preprocessing.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b00bb699-f521-4146-9cf5-905510126ea1Cited by top-tier papers5
- cedar: Optimized and Unified Machine Learning Input Data PipelinesMark Zhao, Emanuel Adamiak, Christos KozyrakisVLDB 2025 · 13 citations
- TensorSocket: Shared Data Loading for Deep Learning TrainingTies Robroek, Neil Kim Nielsen, Pinar TözünSIGMOD 2026 · 4 citations
- Preparation Meets Opportunity: Enhancing Data Preprocessing for ML Training With SenecaOmkar Desai, Ziyang Jiao, Shuyi Pei, Janki Bhimani et al.FAST 2026 · 3 citations
- MinatoLoader: Accelerating Machine Learning Training Through Efficient Data PreprocessingRahma Nouaji, Stella Bitchebe, Ricardo Macedo, Oana BalmauEuroSys 2026 · 3 citations
- MegaScale-Data: Scaling DataLoader for Multisource Large Foundation Model TrainingJuntao Zhao, Qi Lu, Wei Jia, Borui Wan et al.EuroSys 2026
Builds on21
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- RandAugment: Practical Automated Data Augmentation with a Reduced Search SpaceEkin Dogus Cubuk, Barret Zoph, Jonathon Shlens, Quoc LeNeurIPS 2020 · 4,453 citations
- AntMan: Dynamic Scaling on GPU Clusters for Deep LearningWencong Xiao, Shiru Ren, Yong Li, Yang Zhang et al.OSDI 2020 · 260 citations
- Analyzing and Mitigating Data Stalls in DNN TrainingJayashree Mohan, Amar Phanishayee, Ashish Raniwala, Vijay ChidambaramVLDB 2021 · 142 citations
Related papers
- Cachew: Machine Learning Input Data Processing as a ServiceDan Graur, Damien Aymon, Dan Kluser, Tanguy Albrici et al.USENIX ATC 2022 · 43 citations
- FastFlow: Accelerating Deep Learning Model Training with Smart Offloading of Input Data PipelineTaegeon Um, Byungsoo Oh, Byeongchan Seo, Minhyeok Kweun et al.VLDB 2023 · 45 citations
- Enabling Efficient Large-Scale Deep Learning Training with Cache Coherent Disaggregated Memory SystemsZixuan Wang, Joonseop Sim, Euicheol Lim, Jishen ZhaoHPCA 2022 · 9 citations
- PD3: Prefetching Data with DPUs for Disaggregated MemorySidharth Sankhe, Felix Zhang, Umayrah Chonee, Sherman Lim et al.NSDI 2026 · 1 citation
- TrainBox: An Extreme-Scale Neural Network Training Server Architecture by Systematically Balancing OperationsPyeongsu Park, Heetaek Jeong, Jangwoo KimMICRO 2020 · 11 citations
