Efficiently Estimating Mutual Information Between Attributes Across Tables
Aécio S. R. Santos, Flip Korn, Juliana Freire
摘要
Relational data augmentation is a powerful technique for enhancing data analytics and improving machine learning models by incorporating columns from external datasets. However, it is challenging to efficiently discover relevant external tables to join with a given input table. Existing approaches rely on data discovery systems to identify “joinable” tables from external sources, typically based on overlap or containment. However, the sheer number of tables obtained from these systems results in irrelevant joins that need to be performed; this can be computationally expensive or even infeasible in practice. We address this limitation by proposing the use of efficient mutual information (MI) estimation for finding relevant joinable tables. We introduce a new sketching method that enables efficient evaluation of relationship discovery queries by estimating MI without materializing the joins and returning a smaller set of tables that are more likely to be relevant. We also demonstrate the effectiveness of our approach at approximating MI in extensive experiments using synthetic and real-world datasets.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- EDPC: Accelerating Lossless Compression via Lightweight Probability Models and Decoupled Parallel DataflowZeyi Lu, Xiaoxiao Ma, Yujun Huang, Minxiao Chen 等ACM MM 2025 · 被引用 1 次
- Exploiting LLMs for Automatic Hypothesis Assessment via a Logit-Based Calibrated PriorYue Gong, Raul Castro FernandezNeurIPS 2025 · 被引用 1 次
- MosaicJoin: Compact Semantic Sketches for Value-Level Join DiscoveryGrace Fan, Eden Wu, Majid Daliri, Juliana FreireVLDB 2026 · 被引用 1 次
- Matryoshka: Uncovering Relevant Features in Data Lakes to Enhance Machine Learning ApplicationsFedor Turchenko, Runjie Zhang, Binger Chen, Matthias Boehm 等VLDB 2026
它引用的顶会 Paper13
- Correlation Sketches for Approximate Join-Correlation QueriesAécio S. R. Santos, Aline Bessa, Fernando Chirigati, Christopher Musco 等SIGMOD 2021 · 被引用 45 次
- A Sketch-based Index for Correlated Dataset SearchAécio S. R. Santos, Aline Bessa, Christopher Musco, Juliana FreireICDE 2022 · 被引用 31 次
- MATE: Multi-Attribute Table ExtractionMahdi Esmailoghli, Jorge-Arnulfo Quiané-Ruiz, Ziawasch AbedjanVLDB 2022 · 被引用 30 次
- Metam: Goal-Oriented Data DiscoverySainyam Galhotra, Yue Gong, Raul Castro FernandezICDE 2023 · 被引用 28 次
- Towards Benchmarking Feature Type Inference for AutoML PlatformsVraj Shah, Jonathan Lacanlale, Premanand Kumar, Kevin Yang 等SIGMOD 2021 · 被引用 16 次
相关 Paper
- TabSketchFM: Sketch-Based Tabular Representation Learning for Data Discovery Over Data LakesAamod Khatiwada, Harsha Kokel, Ibrahim Abdelaziz, Subhajit Chaudhury 等ICDE 2025 · 被引用 3 次
- Coresets over Multiple Tables for Feature-rich and Data-efficient Machine LearningJiayi Wang, Chengliang Chai, Nan Tang, Jiabin Liu 等VLDB 2023 · 被引用 31 次
- AutoFeat: Transitive Feature Discovery over Join PathsAndra Ionescu, Kiril Vasilev, Florena Buse, Rihan Hai 等ICDE 2024 · 被引用 12 次
- Putting Things into Context: Rich Explanations for Query Answers using Join GraphsChenjie Li, Zhengjie Miao, Qitian Zeng, Boris Glavic 等SIGMOD 2021 · 被引用 16 次
- Distinctiveness Maximization in Datasets AssemblageTingting Wang, Shixun Huang, Zhifeng Bao, J. Shane Culpepper 等WWW 2025 · 被引用 3 次
