Leva: Boosting Machine Learning Performance with Relational Embedding Data Augmentation
Zixuan Zhao, Raul Castro Fernandez
Abstract
In this paper, we present Leva, an end-to-end system that boosts the performance of machine learning tasks over relational data. Leva builds a relational embedding by representing relational data as a graph and then using embedding methods to represent the graph as vectors. The embedding represents information from the entire database, including useful information for the downstream machine learning task. At the same time, some information in the graph will be erroneous, for example, corresponding to incorrect inclusion dependencies. However, we show that the supervision signal from the downstream task filters out information that is not useful. The result is a boost in ML performance. This result means that it is possible for analysts to avoid the time-consuming effort of collecting features across multiple relations-which requires solving a data discovery and integration problem-and instead rely on these techniques to train better-performing models. We demonstrate Leva's performance on different classification and regression datasets and compare it with multiple other baselines.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext aa488901-8ba5-43df-979d-912e905016d1Cited by top-tier papers12
- Semantics-aware Dataset Discovery from Data Lakes with Contextualized Column-based Representation LearningGrace Fan, Jin Wang, Yuliang Li, Dan Zhang et al.VLDB 2023 · 139 citations
- How Large Language Models Will Disrupt Data ManagementRaul Castro Fernandez, Aaron J. Elmore, Michael J. Franklin, Sanjay Krishnan et al.VLDB 2023 · 127 citations
- Metam: Goal-Oriented Data DiscoverySainyam Galhotra, Yue Gong, Raul Castro FernandezICDE 2023 · 28 citations
- DiffPrep: Differentiable Data Preprocessing Pipeline Search for Learning over Tabular DataPeng Li, Zhiyi Chen, Xu Chu, Kexin RongSIGMOD 2023 · 24 citations
- Solo: Data Discovery Using Natural Language Questions Via A Self-Supervised ApproachQiming Wang, Raul Castro FernandezSIGMOD 2024 · 20 citations
Builds on5
- TURL: Table Understanding through Representation LearningXiang Deng, Huan Sun, Alyssa Lees, You Wu et al.VLDB 2021 · 2,406 citations
- Creating Embeddings of Heterogeneous Relational Datasets for Data Integration TasksRiccardo Cappuzzo, Paolo Papotti, Saravanan ThirumuruganathanSIGMOD 2020 · 139 citations
- Correlation Sketches for Approximate Join-Correlation QueriesAécio S. R. Santos, Aline Bessa, Fernando Chirigati, Christopher Musco et al.SIGMOD 2021 · 45 citations
- ARDA: Automatic Relational Data Augmentation for Machine LearningNadiia Chepurko, Ryan Marcus, Emanuel Zgraggen, Raul Castro Fernandez et al.VLDB 2020 · 14 citations
- Learning Over Dirty Data Without CleaningJose Picado, John Davis, Arash Termehchy, Ga Young LeeSIGMOD 2020 · 13 citations
Related papers
- InferDB: In-Database Machine Learning Inference Using IndexesRicardo Salazar-Díaz, Boris Glavic, Tilmann RablVLDB 2024 · 13 citations
- Leveraging Relational Graph Neural Network for Transductive Model EnsembleZhengyu Hu, Jieyu Zhang, Haonan Wang, Siwei Liu et al.KDD 2023 · 8 citations
- NeurIDA: Dynamic Modeling for Effective In-Database AnalyticsLingze Zeng, Shaofeng Cai, Naili Xing, Jiaqi Zhu et al.VLDB 2026
- PerfGuard: Deploying ML-for-Systems without Performance Regressions, Almost!H. M. Sajjad Hossain, Marc T. Friedman, Hiren Patel, Shi Qiao et al.VLDB 2021 · 10 citations
- Graph-based Relation Mining for Context-free Out-of-vocabulary Word Embedding LearningZiran Liang, Yuyin Lu, Hegang Chen, Yanghui RaoACL 2023 · 5 citations
