Leva: Boosting Machine Learning Performance with Relational Embedding Data Augmentation
Zixuan Zhao, Raul Castro Fernandez
摘要
In this paper, we present Leva, an end-to-end system that boosts the performance of machine learning tasks over relational data. Leva builds a relational embedding by representing relational data as a graph and then using embedding methods to represent the graph as vectors. The embedding represents information from the entire database, including useful information for the downstream machine learning task. At the same time, some information in the graph will be erroneous, for example, corresponding to incorrect inclusion dependencies. However, we show that the supervision signal from the downstream task filters out information that is not useful. The result is a boost in ML performance. This result means that it is possible for analysts to avoid the time-consuming effort of collecting features across multiple relations-which requires solving a data discovery and integration problem-and instead rely on these techniques to train better-performing models. We demonstrate Leva's performance on different classification and regression datasets and compare it with multiple other baselines.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper12
- Semantics-aware Dataset Discovery from Data Lakes with Contextualized Column-based Representation LearningGrace Fan, Jin Wang, Yuliang Li, Dan Zhang 等VLDB 2023 · 被引用 139 次
- How Large Language Models Will Disrupt Data ManagementRaul Castro Fernandez, Aaron J. Elmore, Michael J. Franklin, Sanjay Krishnan 等VLDB 2023 · 被引用 127 次
- Metam: Goal-Oriented Data DiscoverySainyam Galhotra, Yue Gong, Raul Castro FernandezICDE 2023 · 被引用 28 次
- DiffPrep: Differentiable Data Preprocessing Pipeline Search for Learning over Tabular DataPeng Li, Zhiyi Chen, Xu Chu, Kexin RongSIGMOD 2023 · 被引用 24 次
- Solo: Data Discovery Using Natural Language Questions Via A Self-Supervised ApproachQiming Wang, Raul Castro FernandezSIGMOD 2024 · 被引用 20 次
它引用的顶会 Paper5
- TURL: Table Understanding through Representation LearningXiang Deng, Huan Sun, Alyssa Lees, You Wu 等VLDB 2021 · 被引用 2,406 次
- Creating Embeddings of Heterogeneous Relational Datasets for Data Integration TasksRiccardo Cappuzzo, Paolo Papotti, Saravanan ThirumuruganathanSIGMOD 2020 · 被引用 139 次
- Correlation Sketches for Approximate Join-Correlation QueriesAécio S. R. Santos, Aline Bessa, Fernando Chirigati, Christopher Musco 等SIGMOD 2021 · 被引用 45 次
- ARDA: Automatic Relational Data Augmentation for Machine LearningNadiia Chepurko, Ryan Marcus, Emanuel Zgraggen, Raul Castro Fernandez 等VLDB 2020 · 被引用 14 次
- Learning Over Dirty Data Without CleaningJose Picado, John Davis, Arash Termehchy, Ga Young LeeSIGMOD 2020 · 被引用 13 次
相关 Paper
- InferDB: In-Database Machine Learning Inference Using IndexesRicardo Salazar-Díaz, Boris Glavic, Tilmann RablVLDB 2024 · 被引用 13 次
- Leveraging Relational Graph Neural Network for Transductive Model EnsembleZhengyu Hu, Jieyu Zhang, Haonan Wang, Siwei Liu 等KDD 2023 · 被引用 8 次
- NeurIDA: Dynamic Modeling for Effective In-Database AnalyticsLingze Zeng, Shaofeng Cai, Naili Xing, Jiaqi Zhu 等VLDB 2026
- PerfGuard: Deploying ML-for-Systems without Performance Regressions, Almost!H. M. Sajjad Hossain, Marc T. Friedman, Hiren Patel, Shi Qiao 等VLDB 2021 · 被引用 10 次
- Graph-based Relation Mining for Context-free Out-of-vocabulary Word Embedding LearningZiran Liang, Yuyin Lu, Hegang Chen, Yanghui RaoACL 2023 · 被引用 5 次
