A Sketch-based Index for Correlated Dataset Search
Aécio S. R. Santos, Aline Bessa, Christopher Musco, Juliana Freire
摘要
Dataset search is emerging as a critical capability in both research and industry: it has spurred many novel applications, ranging from the enrichment of analyses of real-world phenomena to the improvement of machine learning models. Recent research in this field has explored a new class of data-driven queries: queries consist of datasets and retrieve, from a large collection, related datasets. In this paper, we study a specific type of data-driven query that supports relational data augmentation through numerical data relationships: given an input query table, find the top-k tables that are both joinable with it and contain columns that are correlated with a column in the query. We propose a novel hashing scheme that allows the construction of a sketch-based index to support efficient correlated table search. We show that our proposed approach is effective and efficient, and achieves better trade-offs that significantly improve both the ranking accuracy and recall compared to the state-of-the-art solutions.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper14
- Semantics-aware Dataset Discovery from Data Lakes with Contextualized Column-based Representation LearningGrace Fan, Jin Wang, Yuliang Li, Dan Zhang 等VLDB 2023 · 被引用 139 次
- CHORUS: Foundation Models for Unified Data Discovery and ExplorationMoe Kayali, Anton Lykov, Ilias Fountalis, Nikolaos Vasiloglou 等VLDB 2024 · 被引用 33 次
- HyperCalm Sketch: One-Pass Mining Periodic Batches in Data StreamsZirui Liu, Chaozhe Kong, Kaicheng Yang, Tong Yang 等ICDE 2023 · 被引用 16 次
- Saibot: A Differentially Private Data Search PlatformZezhou Huang, Jiaxiang Liu, Daniel Alabi, Raul Castro Fernandez 等VLDB 2023 · 被引用 13 次
- Fainder: A Fast and Accurate Index for Distribution-Aware Dataset SearchLennart Behme, Sainyam Galhotra, Kaustubh Beedkar, Volker MarklVLDB 2024 · 被引用 9 次
它引用的顶会 Paper5
- Finding Related Tables in Data Lakes for Interactive Data ScienceYi Zhang, Zachary G. IvesSIGMOD 2020 · 被引用 98 次
- Correlation Sketches for Approximate Join-Correlation QueriesAécio S. R. Santos, Aline Bessa, Fernando Chirigati, Christopher Musco 等SIGMOD 2021 · 被引用 45 次
- ARDA: Automatic Relational Data Augmentation for Machine LearningNadiia Chepurko, Ryan Marcus, Emanuel Zgraggen, Raul Castro Fernandez 等VLDB 2020 · 被引用 14 次
- Finding the Best of Both Worlds: Faster and More Robust Top-k Document RetrievalOmar Khattab, Mohammad Hammoud, Tamer ElsayedSIGIR 2020 · 被引用 12 次
- An Ecosystem of Applications for Modeling Political ViolenceAline Bessa, Sonia Castelo, Rémi Rampin, Aécio S. R. Santos 等SIGMOD 2021 · 被引用 4 次
相关 Paper
- Efficiently Estimating Mutual Information Between Attributes Across TablesAécio S. R. Santos, Flip Korn, Juliana FreireICDE 2024 · 被引用 2 次
- TabSketchFM: Sketch-Based Tabular Representation Learning for Data Discovery Over Data LakesAamod Khatiwada, Harsha Kokel, Ibrahim Abdelaziz, Subhajit Chaudhury 等ICDE 2025 · 被引用 3 次
- Dataset Discovery in Data LakesAlex Bogatu, Alvaro A. A. Fernandes, Norman W. Paton, Nikolaos KonstantinouICDE 2020 · 被引用 118 次
- Decomposition-Driven Multi-Table Retrieval and Reasoning for Numerical Question AnsweringFeng Luo, Hai Lan, Hui Luo, Zhifeng Bao 等ICDE 2026 · 被引用 1 次
- Putting Things into Context: Rich Explanations for Query Answers using Join GraphsChenjie Li, Zhengjie Miao, Qitian Zeng, Boris Glavic 等SIGMOD 2021 · 被引用 16 次
