LakeBench: A Benchmark for Discovering Joinable and Unionable Tables in Data Lakes
Yuhao Deng, Chengliang Chai, Lei Cao, Qin Yuan, Siyuan Chen, Yanrui Yu, Zhaoze Sun, Junyi Wang, Jiajun Li, Ziqi Cao, Kaisen Jin, Chi Zhang
Abstract
Discovering tables from poorly maintained data lakes is a significant challenge in data management. Two key tasks are identifying joinable and unionable tables, crucial for data integration, analysis, and machine learning. However, there's a lack of a comprehensive benchmark for evaluating existing methods. To address this, we introduce LakeBench, a large-scale table discovery benchmark. It evaluates effectiveness, efficiency, and scalability of table join & union search methods. With over 16 million real tables, LakeBench is 1,600X larger than existing datasets and 100X larger in storage size. It includes synthesized and real queries with ground truth, totaling more than 10 thousand queries - 10X more than used in any existing evaluation. We spent over 7,500 human hours labeling these queries and constructing diverse query categories for thorough evaluation. Our benchmark thoroughly evaluates state-of-the-art table discovery methods, providing insights into their performance and highlighting research opportunities.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext f4ecfdd3-0107-49eb-a0b7-cebbb5420d4cCited by top-tier papers13
- BIRDIE: Natural Language-Driven Table Discovery Using Differentiable Search IndexYuxiang Guo, Zhonghao Hu, Yuren Mao, Baihua Zheng et al.VLDB 2025 · 6 citations
- Qualitative Join Discovery in Data Lakes using ExamplesMir Mahathir Mohammad, El Kindi RezigSIGMOD 2026 · 6 citations
- BLEND: A Unified Data Discovery SystemMahdi Esmailoghli, Christoph Schnell, Renée J. Miller, Ziawasch AbedjanICDE 2025 · 4 citations
- LakeVisage: Towards Scalable, Flexible and Interactive Visualization Recommendation for Data Discovery over Data LakesYihao Hu, Jin Wang, Sajjadur RahmanVLDB 2025 · 3 citations
- TabSketchFM: Sketch-Based Tabular Representation Learning for Data Discovery Over Data LakesAamod Khatiwada, Harsha Kokel, Ibrahim Abdelaziz, Subhajit Chaudhury et al.ICDE 2025 · 3 citations
Builds on17
- MPNet: Masked and Permuted Pre-training for Language UnderstandingKaitao Song, Xu Tan, Tao Qin, Jianfeng Lu et al.NeurIPS 2020 · 1,957 citations
- TaBERT: Pretraining for Joint Understanding of Textual and Tabular DataPengcheng Yin, Graham Neubig, Wen-tau Yih, Sebastian RiedelACL 2020 · 417 citations
- Data Valuation using Reinforcement LearningJinsung Yoon, Sercan Ömer Arik, Tomas PfisterICML 2020 · 236 citations
- Semantics-aware Dataset Discovery from Data Lakes with Contextualized Column-based Representation LearningGrace Fan, Jin Wang, Yuliang Li, Dan Zhang et al.VLDB 2023 · 139 citations
- Dataset Discovery in Data LakesAlex Bogatu, Alvaro A. A. Fernandes, Norman W. Paton, Nikolaos KonstantinouICDE 2020 · 118 citations
Related papers
- Integrating Data Lake TablesAamod Khatiwada, Roee Shraga, Wolfgang Gatterbauer, Renée J. MillerVLDB 2023 · 59 citations
- SANTOS: Relationship-based Semantic Table Union SearchAamod Khatiwada, Grace Fan, Roee Shraga, Zixuan Chen et al.SIGMOD 2023 · 61 citations
- Human-Centered Exploration of Table UnionabilityNina Klimenkova, Sreeram Marimuthu, Roee ShragaVLDB 2026
- Gen-T: Table Reclamation in Data LakesGrace Fan, Roee Shraga, Renée J. MillerICDE 2024 · 5 citations
- Revisiting Single-Table Retrieval: An Open Problem Under 360° Stress TestsChenyu Yang, Ziyu Jiang, Junhao Li, Yuyu Luo et al.ICDE 2026
