UPLIFT: Parallelization Strategies for Feature Transformations in Machine Learning Workloads
Arnab Phani, Lukas Erlbacher, Matthias Boehm
Abstract
Data science pipelines are typically exploratory. An integral task of such pipelines are feature transformations, which transform raw data into numerical matrices or tensors for training or scoring. There exist a wide variety of transformations for different data modalities. These feature transformations incur large computational overhead due to expensive string processing and dictionary creation. Existing ML systems address this overhead by static parallelization schemes and interleaving transformations with model training. These approaches show good performance improvements for simple transformations, but struggle to handle different data characteristics (many features/distinct items) and multi-pass transformations. A key observation is that good parallelization strategies for feature transformations depend on data characteristics. In this paper, we introduce UPLIFT, a framework for P aralle LI zing F eature T ransformations. UPLIFT constructs a fine-grained task graph for a set of transformations, optimizes the plan according to data characteristics, and executes this plan in a cache-conscious manner. We show that the resulting framework is applicable to a wide range of transformations. Furthermore, we propose the FTBench benchmark with transformations and datasets from various domains. On this benchmark, UPLIFT yields speedups of up to 31.6x (9.27x on average) compared to state-of-the-art ML systems.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext abfb6b5e-4e58-41a5-baf1-5facc56524afCited by top-tier papers4
- Automating and Optimizing Data-Centric What-If Analyses on Native Machine Learning PipelinesStefan Grafberger, Paul Groth, Sebastian SchelterSIGMOD 2023 · 18 citations
- Optimizing Data Pipelines for Machine Learning in Feature StoresRui Liu, Kwanghyun Park, Fotis Psallidas, Xiaoyong Zhu et al.VLDB 2023 · 10 citations
- DFLOP: A Data-driven Framework for Multimodal LLM Training Pipeline OptimizationHyeonjun An, Sihyun Kim, Chaerim Lim, Hyunjoon Kim et al.SIGMOD 2026 · 1 citation
- stratum: A System Infrastructure for Massive Agent-Centric ML WorkloadsArnab Phani, Elias Strauss, Sebastian SchelterVLDB 2026
Builds on11
- CleanML: A Study for Evaluating the Impact of Data Cleaning on ML Classification TasksPeng Li, Xi Rao, Jennifer Blase, Yue Zhang et al.ICDE 2021 · 127 citations
- Towards Scalable Dataframe SystemsDevin Petersohn, William W. Ma, Doris Jung Lin Lee, Stephen Macke et al.VLDB 2020 · 109 citations
- Pump Up the Volume: Processing Large Data on GPUs with Fast InterconnectsClemens Lutz, Sebastian Breß, Steffen Zeuch, Tilmann Rabl et al.SIGMOD 2020 · 99 citations
- Cerebro: A Data System for Optimized Deep Learning Model SelectionSupun Nakandala, Yuhao Zhang, Arun KumarVLDB 2020 · 61 citations
- To Partition, or Not to Partition, That is the Join Question in a Real SystemMaximilian Bandle, Jana Giceva, Thomas NeumannSIGMOD 2021 · 43 citations
Related papers
- ArrayPipe: Introducing Job-Array Pipeline Parallelism for High Throughput Model ExplorationHairui Zhao, Hongliang Li, Qi Tian, Jie Wu et al.INFOCOM 2025 · 3 citations
- FastFlow: Accelerating Deep Learning Model Training with Smart Offloading of Input Data PipelineTaegeon Um, Byungsoo Oh, Byeongchan Seo, Minhyeok Kweun et al.VLDB 2023 · 45 citations
- Unity: Accelerating DNN Training Through Joint Optimization of Algebraic Transformations and ParallelizationColin Unger, Zhihao Jia, Wei Wu, Sina Lin et al.OSDI 2022 · 105 citations
- Nimble: Lightweight and Parallel GPU Task Scheduling for Deep LearningWoosuk Kwon, Gyeong-In Yu, Eunji Jeong, Byung-Gon ChunNeurIPS 2020 · 102 citations
- Fine-tuning giant neural networks on commodity hardware with automatic pipeline model parallelismSaar Eliad, Ido Hakimi, Alon De Jagger, Mark Silberstein et al.USENIX ATC 2021 · 24 citations
