Efficiently Estimating Mutual Information Between Attributes Across Tables
Aécio S. R. Santos, Flip Korn, Juliana Freire
Abstract
Relational data augmentation is a powerful technique for enhancing data analytics and improving machine learning models by incorporating columns from external datasets. However, it is challenging to efficiently discover relevant external tables to join with a given input table. Existing approaches rely on data discovery systems to identify “joinable” tables from external sources, typically based on overlap or containment. However, the sheer number of tables obtained from these systems results in irrelevant joins that need to be performed; this can be computationally expensive or even infeasible in practice. We address this limitation by proposing the use of efficient mutual information (MI) estimation for finding relevant joinable tables. We introduce a new sketching method that enables efficient evaluation of relationship discovery queries by estimating MI without materializing the joins and returning a smaller set of tables that are more likely to be relevant. We also demonstrate the effectiveness of our approach at approximating MI in extensive experiments using synthetic and real-world datasets.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e2734ca0-5f14-4b96-8055-a01a24d2ef83Cited by top-tier papers4
- EDPC: Accelerating Lossless Compression via Lightweight Probability Models and Decoupled Parallel DataflowZeyi Lu, Xiaoxiao Ma, Yujun Huang, Minxiao Chen et al.ACM MM 2025 · 1 citation
- Exploiting LLMs for Automatic Hypothesis Assessment via a Logit-Based Calibrated PriorYue Gong, Raul Castro FernandezNeurIPS 2025 · 1 citation
- MosaicJoin: Compact Semantic Sketches for Value-Level Join DiscoveryGrace Fan, Eden Wu, Majid Daliri, Juliana FreireVLDB 2026 · 1 citation
- Matryoshka: Uncovering Relevant Features in Data Lakes to Enhance Machine Learning ApplicationsFedor Turchenko, Runjie Zhang, Binger Chen, Matthias Boehm et al.VLDB 2026
Builds on13
- Correlation Sketches for Approximate Join-Correlation QueriesAécio S. R. Santos, Aline Bessa, Fernando Chirigati, Christopher Musco et al.SIGMOD 2021 · 45 citations
- A Sketch-based Index for Correlated Dataset SearchAécio S. R. Santos, Aline Bessa, Christopher Musco, Juliana FreireICDE 2022 · 31 citations
- MATE: Multi-Attribute Table ExtractionMahdi Esmailoghli, Jorge-Arnulfo Quiané-Ruiz, Ziawasch AbedjanVLDB 2022 · 30 citations
- Metam: Goal-Oriented Data DiscoverySainyam Galhotra, Yue Gong, Raul Castro FernandezICDE 2023 · 28 citations
- Towards Benchmarking Feature Type Inference for AutoML PlatformsVraj Shah, Jonathan Lacanlale, Premanand Kumar, Kevin Yang et al.SIGMOD 2021 · 16 citations
Related papers
- TabSketchFM: Sketch-Based Tabular Representation Learning for Data Discovery Over Data LakesAamod Khatiwada, Harsha Kokel, Ibrahim Abdelaziz, Subhajit Chaudhury et al.ICDE 2025 · 3 citations
- Coresets over Multiple Tables for Feature-rich and Data-efficient Machine LearningJiayi Wang, Chengliang Chai, Nan Tang, Jiabin Liu et al.VLDB 2023 · 31 citations
- AutoFeat: Transitive Feature Discovery over Join PathsAndra Ionescu, Kiril Vasilev, Florena Buse, Rihan Hai et al.ICDE 2024 · 12 citations
- Putting Things into Context: Rich Explanations for Query Answers using Join GraphsChenjie Li, Zhengjie Miao, Qitian Zeng, Boris Glavic et al.SIGMOD 2021 · 16 citations
- Distinctiveness Maximization in Datasets AssemblageTingting Wang, Shixun Huang, Zhifeng Bao, J. Shane Culpepper et al.WWW 2025 · 3 citations
