A Sketch-based Index for Correlated Dataset Search
Aécio S. R. Santos, Aline Bessa, Christopher Musco, Juliana Freire
Abstract
Dataset search is emerging as a critical capability in both research and industry: it has spurred many novel applications, ranging from the enrichment of analyses of real-world phenomena to the improvement of machine learning models. Recent research in this field has explored a new class of data-driven queries: queries consist of datasets and retrieve, from a large collection, related datasets. In this paper, we study a specific type of data-driven query that supports relational data augmentation through numerical data relationships: given an input query table, find the top-k tables that are both joinable with it and contain columns that are correlated with a column in the query. We propose a novel hashing scheme that allows the construction of a sketch-based index to support efficient correlated table search. We show that our proposed approach is effective and efficient, and achieves better trade-offs that significantly improve both the ranking accuracy and recall compared to the state-of-the-art solutions.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers14
- Semantics-aware Dataset Discovery from Data Lakes with Contextualized Column-based Representation LearningGrace Fan, Jin Wang, Yuliang Li, Dan Zhang et al.VLDB 2023 · 139 citations
- CHORUS: Foundation Models for Unified Data Discovery and ExplorationMoe Kayali, Anton Lykov, Ilias Fountalis, Nikolaos Vasiloglou et al.VLDB 2024 · 33 citations
- HyperCalm Sketch: One-Pass Mining Periodic Batches in Data StreamsZirui Liu, Chaozhe Kong, Kaicheng Yang, Tong Yang et al.ICDE 2023 · 16 citations
- Saibot: A Differentially Private Data Search PlatformZezhou Huang, Jiaxiang Liu, Daniel Alabi, Raul Castro Fernandez et al.VLDB 2023 · 13 citations
- Fainder: A Fast and Accurate Index for Distribution-Aware Dataset SearchLennart Behme, Sainyam Galhotra, Kaustubh Beedkar, Volker MarklVLDB 2024 · 9 citations
Builds on5
- Finding Related Tables in Data Lakes for Interactive Data ScienceYi Zhang, Zachary G. IvesSIGMOD 2020 · 98 citations
- Correlation Sketches for Approximate Join-Correlation QueriesAécio S. R. Santos, Aline Bessa, Fernando Chirigati, Christopher Musco et al.SIGMOD 2021 · 45 citations
- ARDA: Automatic Relational Data Augmentation for Machine LearningNadiia Chepurko, Ryan Marcus, Emanuel Zgraggen, Raul Castro Fernandez et al.VLDB 2020 · 14 citations
- Finding the Best of Both Worlds: Faster and More Robust Top-k Document RetrievalOmar Khattab, Mohammad Hammoud, Tamer ElsayedSIGIR 2020 · 12 citations
- An Ecosystem of Applications for Modeling Political ViolenceAline Bessa, Sonia Castelo, Rémi Rampin, Aécio S. R. Santos et al.SIGMOD 2021 · 4 citations
Related papers
- Efficiently Estimating Mutual Information Between Attributes Across TablesAécio S. R. Santos, Flip Korn, Juliana FreireICDE 2024 · 2 citations
- TabSketchFM: Sketch-Based Tabular Representation Learning for Data Discovery Over Data LakesAamod Khatiwada, Harsha Kokel, Ibrahim Abdelaziz, Subhajit Chaudhury et al.ICDE 2025 · 3 citations
- Dataset Discovery in Data LakesAlex Bogatu, Alvaro A. A. Fernandes, Norman W. Paton, Nikolaos KonstantinouICDE 2020 · 118 citations
- Decomposition-Driven Multi-Table Retrieval and Reasoning for Numerical Question AnsweringFeng Luo, Hai Lan, Hui Luo, Zhifeng Bao et al.ICDE 2026 · 1 citation
- Putting Things into Context: Rich Explanations for Query Answers using Join GraphsChenjie Li, Zhengjie Miao, Qitian Zeng, Boris Glavic et al.SIGMOD 2021 · 16 citations
