Semantics-aware Dataset Discovery from Data Lakes with Contextualized Column-based Representation Learning
Grace Fan, Jin Wang, Yuliang Li, Dan Zhang, Renée J. Miller
摘要
Dataset discovery from data lakes is essential in many real application scenarios. In this paper, we propose Starmie, an end-to-end framework for dataset discovery from data lakes (with table union search as the main use case). Our proposed framework features a contrastive learning method to train column encoders from pre-trained language models in a fully unsupervised manner. The column encoder of Starmie captures the rich contextual semantic information within tables by leveraging a contrastive multi-column pre-training strategy. We utilize the cosine similarity between column embedding vectors as the column unionability score and propose a filter-and-verification framework that allows exploring a variety of design choices to compute the unionability score between two tables accordingly. Empirical results on real table benchmarks show that Starmie outperforms the best-known solutions in the effectiveness of table union search by 6.8 in MAP and recall. Moreover, Starmie is the first to employ the HNSW (Hierarchical Navigable Small World) index to accelerate query processing of table union search which provides a 3,000X performance gain over the linear scan baseline and a 400X performance gain over an LSH index (the state-of-the-art solution for data lake indexing).
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper34
- Integrating Data Lake TablesAamod Khatiwada, Roee Shraga, Wolfgang Gatterbauer, Renée J. MillerVLDB 2023 · 被引用 59 次
- LakeBench: A Benchmark for Discovering Joinable and Unionable Tables in Data LakesYuhao Deng, Chengliang Chai, Lei Cao, Qin Yuan 等VLDB 2024 · 被引用 36 次
- Magneto: Combining Small and Large Language Models for Schema MatchingYurong Liu, Eduardo H. M. Pena, Aécio S. R. Santos, Eden Wu 等VLDB 2025 · 被引用 32 次
- Watchog: A Light-weight Contrastive Learning based Framework for Column AnnotationZhengjie Miao, Jin WangSIGMOD 2024 · 被引用 14 次
- R2D2: Reducing Redundancy and Duplication in Data LakesRaunak Shah, Koyel Mukherjee, Atharv Tyagi, Sai Keerthana Karnam 等SIGMOD 2024 · 被引用 8 次
它引用的顶会 Paper15
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 被引用 24,064 次
- TURL: Table Understanding through Representation LearningXiang Deng, Huan Sun, Alyssa Lees, You Wu 等VLDB 2021 · 被引用 2,406 次
- Deep Entity Matching with Pre-Trained Language ModelsYuliang Li, Jinfeng Li, Yoshihiko Suhara, AnHai Doan 等VLDB 2021 · 被引用 484 次
- TaBERT: Pretraining for Joint Understanding of Textual and Tabular DataPengcheng Yin, Graham Neubig, Wen-tau Yih, Sebastian RiedelACL 2020 · 被引用 417 次
- Creating Embeddings of Heterogeneous Relational Datasets for Data Integration TasksRiccardo Cappuzzo, Paolo Papotti, Saravanan ThirumuruganathanSIGMOD 2020 · 被引用 139 次
相关 Paper
- Human-Centered Exploration of Table UnionabilityNina Klimenkova, Sreeram Marimuthu, Roee ShragaVLDB 2026
- SANTOS: Relationship-based Semantic Table Union SearchAamod Khatiwada, Grace Fan, Roee Shraga, Zixuan Chen 等SIGMOD 2023 · 被引用 61 次
- DeepJoin: Joinable Table Discovery with Pre-trained Language ModelsYuyang Dong, Chuan Xiao, Takuma Nozawa, Masafumi Enomoto 等VLDB 2023 · 被引用 53 次
- Dataset Discovery in Data LakesAlex Bogatu, Alvaro A. A. Fernandes, Norman W. Paton, Nikolaos KonstantinouICDE 2020 · 被引用 118 次
- LIFTus: An Adaptive Multi-Aspect Column Representation Learning for Table Union SearchErmu Qiu, Jun Gao, Yaofeng Tu, Jingru YangICDE 2025
