Qualitative Join Discovery in Data Lakes using Examples
Mir Mahathir Mohammad, El Kindi Rezig
Abstract
Finding relevant datasets is critical in any data pipeline but becomes challenging when data lacks schemas or metadata, as in data lakes. This makes it hard to identify the joins needed to produce the desired dataset. In query-by-example (QbE) join discovery, users provide a query table with a few example values, aiming to find joins from data lake tables that produce datasets containing those examples. Current QbE methods rely only on syntactic similarity, while semantic join discovery methods do not support QbE interfaces that work with limited example values. Moreover, existing QbE join path discovery methods (1) assume that the matching tables are directly joinable with each other, whereas in practice, a join path might contain intermediate tables that don't match the query table; and (2) do not ensure that the example tuples are contained in the returned joined table. We propose SemDisc , an end-to-end join discovery system that provides (1) discovery of hybrid join paths using both equi-join and semantic joins across data lake tables, (2) produces join paths that may include intermediate tables that do not overlap with the query tables but are needed to build high-quality joins, and (3) ensures the returned tuples are semantically similar to the ones in the provided examples. SemDisc supports efficient querying of joinable tables using an index that keeps track of high-quality join paths. Our evaluation across diverse workloads and datasets shows that SemDisc yields an average precision of over 0.86 in finding the correct join paths across various benchmarks, which is more than a 3x improvement over state-of-the-art join discovery methods.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on30
- TURL: Table Understanding through Representation LearningXiang Deng, Huan Sun, Alyssa Lees, You Wu et al.VLDB 2021 · 2,406 citations
- Deep Entity Matching with Pre-Trained Language ModelsYuliang Li, Jinfeng Li, Yoshihiko Suhara, AnHai Doan et al.VLDB 2021 · 484 citations
- TaBERT: Pretraining for Joint Understanding of Textual and Tabular DataPengcheng Yin, Graham Neubig, Wen-tau Yih, Sebastian RiedelACL 2020 · 417 citations
- An End-to-End Learning-based Cost EstimatorJi Sun, Guoliang LiVLDB 2020 · 251 citations
- Are We Ready For Learned Cardinality Estimation?Xiaoying Wang, Changbo Qu, Weiyuan Wu, Jiannan Wang et al.VLDB 2021 · 156 citations
Related papers
- MosaicJoin: Compact Semantic Sketches for Value-Level Join DiscoveryGrace Fan, Eden Wu, Majid Daliri, Juliana FreireVLDB 2026 · 1 citation
- DeepJoin: Joinable Table Discovery with Pre-trained Language ModelsYuyang Dong, Chuan Xiao, Takuma Nozawa, Masafumi Enomoto et al.VLDB 2023 · 53 citations
- Efficient Joinable Table Discovery in Data Lakes: A High-Dimensional Similarity-Based ApproachYuyang Dong, Kunihiro Takeoka, Chuan Xiao, Masafumi OyamadaICDE 2021 · 78 citations
- Dataset Discovery in Data LakesAlex Bogatu, Alvaro A. A. Fernandes, Norman W. Paton, Nikolaos KonstantinouICDE 2020 · 118 citations
- Integrating Data Lake TablesAamod Khatiwada, Roee Shraga, Wolfgang Gatterbauer, Renée J. MillerVLDB 2023 · 59 citations
