Discovering Related Data At Scale
Sagar Bharadwaj, Praveen Gupta, Ranjita Bhagwan, Saikat Guha
摘要
Analysts frequently require data from multiple sources for their tasks, but finding these sources is challenging in exabyte-scale data lakes. In this paper, we address this problem for our enterprise's data lake by using machine-learning to identify related data sources. Leveraging queries made to the data lake over a month, we build a relevance model that determines whether two columns across two data streams are related or not. We then use the model to find relations at scale across tens of millions of column-pairs and thereafter construct a data relationship graph in a scalable fashion, processing a data lake that has 4.5 Petabytes of data in approximately 80 minutes. Using manually labeled datasets as ground-truth, we show that our techniques show improvements of at least 23% when compared to state-of-the-art methods.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- DeepJoin: Joinable Table Discovery with Pre-trained Language ModelsYuyang Dong, Chuan Xiao, Takuma Nozawa, Masafumi Enomoto 等VLDB 2023 · 被引用 53 次
- Fainder: A Fast and Accurate Index for Distribution-Aware Dataset SearchLennart Behme, Sainyam Galhotra, Kaustubh Beedkar, Volker MarklVLDB 2024 · 被引用 9 次
- R2D2: Reducing Redundancy and Duplication in Data LakesRaunak Shah, Koyel Mukherjee, Atharv Tyagi, Sai Keerthana Karnam 等SIGMOD 2024 · 被引用 8 次
- BLEND: A Unified Data Discovery SystemMahdi Esmailoghli, Christoph Schnell, Renée J. Miller, Ziawasch AbedjanICDE 2025 · 被引用 4 次
- OmniMatch: Joinability Discovery in Data ProductsChristos Koutras, Jiani Zhang, Xiao Qin, Chuan Lei 等VLDB 2025 · 被引用 3 次
它引用的顶会 Paper1
相关 Paper
- AutoFeat: Transitive Feature Discovery over Join PathsAndra Ionescu, Kiril Vasilev, Florena Buse, Rihan Hai 等ICDE 2024 · 被引用 12 次
- Dataset Discovery in Data LakesAlex Bogatu, Alvaro A. A. Fernandes, Norman W. Paton, Nikolaos KonstantinouICDE 2020 · 被引用 118 次
- Searching Data Lakes for Nested and Joined DataYi Zhang, Peter Chen, Zack IvesVLDB 2024
- An Effective Framework for Enhancing Query Answering in a Heterogeneous Data LakeQin Yuan, Ye Yuan, Zhenyu Wen, He Wang 等SIGIR 2023 · 被引用 7 次
- Cross Modal Data Discovery over Structured and Unstructured Data LakesMohamed Y. Eltabakh, Mayuresh Kunjir, Ahmed K. Elmagarmid, Mohammad Shahmeer AhmadVLDB 2023 · 被引用 11 次
