Matryoshka: Uncovering Relevant Features in Data Lakes to Enhance Machine Learning Applications
Fedor Turchenko, Runjie Zhang, Binger Chen, Matthias Boehm, Babak Salimi, Amir Shaikhha, Ziawasch Abedjan
摘要
Data coverage and completeness are crucial prerequisites for machine learning (ML) applications. Data discovery allows retrieving relevant features from unstructured data lakes to augment the training data. However, existing methods either target only individual ML tasks, decouple table retrieval from feature selection, or neglect inter-feature relationships, resulting in redundant, high-dimensional data. In this paper, we introduce Matryoshka, a holistic system that discovers and selects relevant features from data lakes for downstream ML applications. Matryoshka replaces costly join materialization with compact Gram matrix sketches and performs feature selection via incrementally fitted linear proxy models with correlation-based pruning, yielding redundancy-free feature sets for regression and classification. Matryoshka improves downstream prediction quality by 19.3% on average while achieving the lowest geometric mean runtime and up to 120x faster execution on join-intensive workloads.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper14
- Large Language Models for Automated Data Science: Introducing CAAFE for Context-Aware Automated Feature EngineeringNoah Hollmann, Samuel Müller, Frank HutterNeurIPS 2023 · 被引用 210 次
- Semantics-aware Dataset Discovery from Data Lakes with Contextualized Column-based Representation LearningGrace Fan, Jin Wang, Yuliang Li, Dan Zhang 等VLDB 2023 · 被引用 139 次
- DeepJoin: Joinable Table Discovery with Pre-trained Language ModelsYuyang Dong, Chuan Xiao, Takuma Nozawa, Masafumi Enomoto 等VLDB 2023 · 被引用 53 次
- GitTables: A Large-Scale Corpus of Relational TablesMadelon Hulsebos, Çagatay Demiralp, Paul GrothSIGMOD 2023 · 被引用 42 次
- LakeBench: A Benchmark for Discovering Joinable and Unionable Tables in Data LakesYuhao Deng, Chengliang Chai, Lei Cao, Qin Yuan 等VLDB 2024 · 被引用 36 次
相关 Paper
- AutoFeat: Transitive Feature Discovery over Join PathsAndra Ionescu, Kiril Vasilev, Florena Buse, Rihan Hai 等ICDE 2024 · 被引用 12 次
- Dataset Discovery in Data LakesAlex Bogatu, Alvaro A. A. Fernandes, Norman W. Paton, Nikolaos KonstantinouICDE 2020 · 被引用 118 次
- TabSketchFM: Sketch-Based Tabular Representation Learning for Data Discovery Over Data LakesAamod Khatiwada, Harsha Kokel, Ibrahim Abdelaziz, Subhajit Chaudhury 等ICDE 2025 · 被引用 3 次
- Searching Data Lakes for Nested and Joined DataYi Zhang, Peter Chen, Zack IvesVLDB 2024
- Integrating Data Lake TablesAamod Khatiwada, Roee Shraga, Wolfgang Gatterbauer, Renée J. MillerVLDB 2023 · 被引用 59 次
