Optimizing Data Pipelines for Machine Learning in Feature Stores
Rui Liu, Kwanghyun Park, Fotis Psallidas, Xiaoyong Zhu, Jinghui Mo, Rathijit Sen, Matteo Interlandi, Konstantinos Karanasos, Yuanyuan Tian, Jesús Camacho-Rodríguez
摘要
Data pipelines (i.e., converting raw data to features) are critical for machine learning (ML) models, yet their development and management is time-consuming. Feature stores have recently emerged as a new "DBMS-for-ML" with the premise of enabling data scientists and engineers to define and manage their data pipelines. While current feature stores fulfill their promise from a functionality perspective, they are resource-hungry---with ample opportunities for implementing database-style optimizations to enhance their performance. In this paper, we propose a novel set of optimizations specifically targeted for point-in-time join, which is a critical operation in data pipelines. We implement these optimizations on top of Feathr: a widely-used feature store, and evaluate them on use cases from both the TPCx-AI benchmark and real-world online retail scenarios. Our thorough experimental analysis shows that our optimizations can accelerate data pipelines by up to 3× over state-of-the-art baselines.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- DFLOP: A Data-driven Framework for Multimodal LLM Training Pipeline OptimizationHyeonjun An, Sihyun Kim, Chaerim Lim, Hyunjoon Kim 等SIGMOD 2026 · 被引用 1 次
- TPCx-AI under the Microscope: A Benchmarking Debt AnalysisIlin Tolovski, Philipp Hildebrandt, Khuzaima Daudjee, Tilmann RablVLDB 2026
- CAPS: Cost-Aware ML Pipeline SelectionAntonios Kontaxakis, Dimitris Sacharidis, Alberto Abelló, Sergi Nadal 等VLDB 2026
它引用的顶会 Paper6
- Automatic View Generation with Deep Learning and Reinforcement LearningHaitao Yuan, Guoliang Li, Ling Feng, Ji Sun 等ICDE 2020 · 被引用 66 次
- KLL±: Approximate Quantile Sketches over Dynamic DatasetsFuheng Zhao, Sujaya Maiyya, Ryan Weiner, Divy Agrawal 等VLDB 2021 · 被引用 36 次
- Optimizing An In-memory Database System For AI-powered On-line Decision Augmentation Using Persistent MemoryCheng Chen, Jun Yang, Mian Lu, Taize Wang 等VLDB 2021 · 被引用 19 次
- UPLIFT: Parallelization Strategies for Feature Transformations in Machine Learning WorkloadsArnab Phani, Lukas Erlbacher, Matthias BoehmVLDB 2022 · 被引用 13 次
- Materialization and Reuse Optimizations for Production Data Science PipelinesBehrouz Derakhshan, Alireza Rezaei Mahdiraji, Zoi Kaoudi, Tilmann Rabl 等SIGMOD 2022 · 被引用 12 次
相关 Paper
- RALF: Accuracy-Aware Scheduling for Feature Store MaintenanceSarah Wooders, Xiangxi Mo, Amit Narang, Kevin Lin 等VLDB 2024 · 被引用 6 次
- Featpilot: Automatic Feature Augmentation on Tabular DataJiaming Liang, Chuan Lei, Xiao Qin, Jiani Zhang 等ICDE 2025
- BLEND: A Unified Data Discovery SystemMahdi Esmailoghli, Christoph Schnell, Renée J. Miller, Ziawasch AbedjanICDE 2025 · 被引用 4 次
- AutoFeat: Transitive Feature Discovery over Join PathsAndra Ionescu, Kiril Vasilev, Florena Buse, Rihan Hai 等ICDE 2024 · 被引用 12 次
- DB4ML - An In-Memory Database Kernel with Machine Learning SupportMatthias Jasny, Tobias Ziegler, Tim Kraska, Uwe Röhm 等SIGMOD 2020 · 被引用 27 次
