Optimizing Data Pipelines for Machine Learning in Feature Stores
Rui Liu, Kwanghyun Park, Fotis Psallidas, Xiaoyong Zhu, Jinghui Mo, Rathijit Sen, Matteo Interlandi, Konstantinos Karanasos, Yuanyuan Tian, Jesús Camacho-Rodríguez
Abstract
Data pipelines (i.e., converting raw data to features) are critical for machine learning (ML) models, yet their development and management is time-consuming. Feature stores have recently emerged as a new "DBMS-for-ML" with the premise of enabling data scientists and engineers to define and manage their data pipelines. While current feature stores fulfill their promise from a functionality perspective, they are resource-hungry---with ample opportunities for implementing database-style optimizations to enhance their performance. In this paper, we propose a novel set of optimizations specifically targeted for point-in-time join, which is a critical operation in data pipelines. We implement these optimizations on top of Feathr: a widely-used feature store, and evaluate them on use cases from both the TPCx-AI benchmark and real-world online retail scenarios. Our thorough experimental analysis shows that our optimizations can accelerate data pipelines by up to 3× over state-of-the-art baselines.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 34d629db-5977-4837-843b-1e95e141c78dCited by top-tier papers3
- DFLOP: A Data-driven Framework for Multimodal LLM Training Pipeline OptimizationHyeonjun An, Sihyun Kim, Chaerim Lim, Hyunjoon Kim et al.SIGMOD 2026 · 1 citation
- TPCx-AI under the Microscope: A Benchmarking Debt AnalysisIlin Tolovski, Philipp Hildebrandt, Khuzaima Daudjee, Tilmann RablVLDB 2026
- CAPS: Cost-Aware ML Pipeline SelectionAntonios Kontaxakis, Dimitris Sacharidis, Alberto Abelló, Sergi Nadal et al.VLDB 2026
Builds on6
- Automatic View Generation with Deep Learning and Reinforcement LearningHaitao Yuan, Guoliang Li, Ling Feng, Ji Sun et al.ICDE 2020 · 66 citations
- KLL±: Approximate Quantile Sketches over Dynamic DatasetsFuheng Zhao, Sujaya Maiyya, Ryan Weiner, Divy Agrawal et al.VLDB 2021 · 36 citations
- Optimizing An In-memory Database System For AI-powered On-line Decision Augmentation Using Persistent MemoryCheng Chen, Jun Yang, Mian Lu, Taize Wang et al.VLDB 2021 · 19 citations
- UPLIFT: Parallelization Strategies for Feature Transformations in Machine Learning WorkloadsArnab Phani, Lukas Erlbacher, Matthias BoehmVLDB 2022 · 13 citations
- Materialization and Reuse Optimizations for Production Data Science PipelinesBehrouz Derakhshan, Alireza Rezaei Mahdiraji, Zoi Kaoudi, Tilmann Rabl et al.SIGMOD 2022 · 12 citations
Related papers
- RALF: Accuracy-Aware Scheduling for Feature Store MaintenanceSarah Wooders, Xiangxi Mo, Amit Narang, Kevin Lin et al.VLDB 2024 · 6 citations
- Featpilot: Automatic Feature Augmentation on Tabular DataJiaming Liang, Chuan Lei, Xiao Qin, Jiani Zhang et al.ICDE 2025
- BLEND: A Unified Data Discovery SystemMahdi Esmailoghli, Christoph Schnell, Renée J. Miller, Ziawasch AbedjanICDE 2025 · 4 citations
- AutoFeat: Transitive Feature Discovery over Join PathsAndra Ionescu, Kiril Vasilev, Florena Buse, Rihan Hai et al.ICDE 2024 · 12 citations
- DB4ML - An In-Memory Database Kernel with Machine Learning SupportMatthias Jasny, Tobias Ziegler, Tim Kraska, Uwe Röhm et al.SIGMOD 2020 · 27 citations
