BLEND: A Unified Data Discovery System
Mahdi Esmailoghli, Christoph Schnell, Renée J. Miller, Ziawasch Abedjan
摘要
Most research on data discovery has so far focused on improving individual discovery operators such as join, correlation, or union discovery. However, in practice, a combination of these techniques and their corresponding indexes may be necessary to support arbitrary discovery tasks. We propose BLEND, a comprehensive data discovery system that supports existing operators and enables their flexible pipelining. BLEND is based on a set of lower-level operators that serve as fundamental building blocks for more complex and sophisticated user tasks. To reduce the execution runtime of discovery pipelines, we propose a unified index structure and a rule- and cost-based optimizer that rewrites SQL statements into low-level operators when possible. We show the superior flexibility and efficiency of our system compared to ad-hoc discovery pipelines and stand-alone solutions.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Qualitative Join Discovery in Data Lakes using ExamplesMir Mahathir Mohammad, El Kindi RezigSIGMOD 2026 · 被引用 6 次
- Rethinking Dataset Discovery with DataScoutRachel Lin, Bhavya Chopra, Wenjing Lin, Shreya Shankar 等UIST 2025 · 被引用 2 次
- FedAugment: Table Augmentation Search over Decentralized Data RepositoriesLennart Behme, Emil Badura, Leonard Geißler, Matthias Böhm 等VLDB 2026
- Matryoshka: Uncovering Relevant Features in Data Lakes to Enhance Machine Learning ApplicationsFedor Turchenko, Runjie Zhang, Binger Chen, Matthias Boehm 等VLDB 2026
它引用的顶会 Paper16
- Semantics-aware Dataset Discovery from Data Lakes with Contextualized Column-based Representation LearningGrace Fan, Jin Wang, Yuliang Li, Dan Zhang 等VLDB 2023 · 被引用 139 次
- Dataset Discovery in Data LakesAlex Bogatu, Alvaro A. A. Fernandes, Norman W. Paton, Nikolaos KonstantinouICDE 2020 · 被引用 118 次
- Finding Related Tables in Data Lakes for Interactive Data ScienceYi Zhang, Zachary G. IvesSIGMOD 2020 · 被引用 98 次
- Efficient Joinable Table Discovery in Data Lakes: A High-Dimensional Similarity-Based ApproachYuyang Dong, Kunihiro Takeoka, Chuan Xiao, Masafumi OyamadaICDE 2021 · 被引用 78 次
- SANTOS: Relationship-based Semantic Table Union SearchAamod Khatiwada, Grace Fan, Roee Shraga, Zixuan Chen 等SIGMOD 2023 · 被引用 61 次
相关 Paper
- Predicate Pushdown for Data Science PipelinesCong Yan, Yin Lin, Yeye HeSIGMOD 2023 · 被引用 15 次
- Machine Learning Inference Pipeline Execution Using Pure SQL Based on Operator FusionQingfeng Pan, Jiahe Zhi, Chenyang Zhang, Chen Xu 等ICDE 2025 · 被引用 2 次
- Cross Modal Data Discovery over Structured and Unstructured Data LakesMohamed Y. Eltabakh, Mayuresh Kunjir, Ahmed K. Elmagarmid, Mohammad Shahmeer AhmadVLDB 2023 · 被引用 11 次
- Integrating Data Lake TablesAamod Khatiwada, Roee Shraga, Wolfgang Gatterbauer, Renée J. MillerVLDB 2023 · 被引用 59 次
- Generalized Sub-Query Fusion for Eliminating Redundant I/O from Big-Data QueriesPartho Sarthi, Kaushik Rajan, Akash Lal, Abhishek Modi 等OSDI 2020 · 被引用 5 次
