Cross Modal Data Discovery over Structured and Unstructured Data Lakes
Mohamed Y. Eltabakh, Mayuresh Kunjir, Ahmed K. Elmagarmid, Mohammad Shahmeer Ahmad
Abstract
Organizations are collecting increasingly large amounts of data for data-driven decision making. These data are often dumped into a centralized repository, e.g., a data lake, consisting of thousands of structured and unstructured datasets. Perversely, such mixture of datasets makes the problem of discovering elements (e.g., tables or documents) that are relevant to a user's query or an analytical task very challenging. Despite the recent efforts in data discovery, the problem remains widely open especially in the two fronts of (1) discovering relationships and relatedness across structured and unstructured datasets-where existing techniques suffer from either scalability, being customized for a specific problem type (e.g., entity matching or data integration), or demolishing the structural properties on its way, and (2) developing a holistic system for integrating various similarity measurements and sketches in an effective way to boost the discovery accuracy. In this paper, we propose a new data discovery system, named CMDL, for addressing these two limitations. CMDL supports the data discovery process over both structured and unstructured data while retaining the structural properties of tables. As a result, CMDL is the only system to date that empowers end-users to seamlessly pipeline the discovery tasks across the two modalities. We propose a novel multi-modal embedding representation that captures the similarities between text documents and tabular columns. The model training relies on labeled datasets generated though weak supervision, and thus the system is domain agnostic and easily generalizable. We evaluate CMDL on three real-world data lakes with diverse applications and show that our system is significantly more effective for cross-modality discovery compared to the search-based baseline techniques. Moreover, CMDL is more accurate and robust to different data types and distributions compared to the state-ofthe-art systems that are limited to only the structured datasets.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 8f7632c1-ede8-49c3-ae12-8a1c1260deebCited by top-tier papers3
- Cardinality Estimation for Similarity Search on High-Dimensional Data Objects: The Impact of Reference ObjectsHai Lan, Shixun Huang, Zhifeng Bao, Renata Borovica-GajicVLDB 2025 · 7 citations
- QueryArtisan: Generating Data Manipulation Codes for Ad-hoc Analysis in Data LakesXiu Tang, Wenhao Liu, Sai Wu, Chang Yao et al.VLDB 2025 · 4 citations
- FedAugment: Table Augmentation Search over Decentralized Data RepositoriesLennart Behme, Emil Badura, Leonard Geißler, Matthias Böhm et al.VLDB 2026
Builds on8
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- TURL: Table Understanding through Representation LearningXiang Deng, Huan Sun, Alyssa Lees, You Wu et al.VLDB 2021 · 2,406 citations
- Dataset Discovery in Data LakesAlex Bogatu, Alvaro A. A. Fernandes, Norman W. Paton, Nikolaos KonstantinouICDE 2020 · 118 citations
- Efficient Joinable Table Discovery in Data Lakes: A High-Dimensional Similarity-Based ApproachYuyang Dong, Kunihiro Takeoka, Chuan Xiao, Masafumi OyamadaICDE 2021 · 78 citations
- Data-Driven Domain Discovery for Structured DatasetsMasayo Ota, Heiko Mueller, Juliana Freire, Divesh SrivastavaVLDB 2020 · 38 citations
Related papers
- DeepJoin: Joinable Table Discovery with Pre-trained Language ModelsYuyang Dong, Chuan Xiao, Takuma Nozawa, Masafumi Enomoto et al.VLDB 2023 · 53 citations
- TabSketchFM: Sketch-Based Tabular Representation Learning for Data Discovery Over Data LakesAamod Khatiwada, Harsha Kokel, Ibrahim Abdelaziz, Subhajit Chaudhury et al.ICDE 2025 · 3 citations
- CrossEM: A Prompt Tuning Framework for Cross-Modal Entity MatchingQin Yuan, Ye Yuan, Zhenyu Wen, Chi Chen et al.ICDE 2025 · 1 citation
- SANTOS: Relationship-based Semantic Table Union SearchAamod Khatiwada, Grace Fan, Roee Shraga, Zixuan Chen et al.SIGMOD 2023 · 61 citations
- Integrating Data Lake TablesAamod Khatiwada, Roee Shraga, Wolfgang Gatterbauer, Renée J. MillerVLDB 2023 · 59 citations
