WikiDBGraph: A Data Management Benchmark Suite for Collaborative Learning Over Database Silos
Zhaomin Wu, Ziyang Wang, Bingsheng He
Abstract
Relational databases are often fragmented across organizations, creating data silos that hinder distributed data management and mining. Collaborative learning (CL) -- techniques that enable multiple parties to train models jointly without sharing raw data -- offers a principled approach to this challenge. However, existing CL frameworks (e.g., federated and split learning) remain limited in real-world deployments. Current CL benchmarks and algorithms primarily target the learning step under assumptions of isolated, aligned, and joinable databases, and they typically neglect the end-to-end data management pipeline, especially preprocessing steps such as table joins and data alignment. In contrast, our analysis of the real-world corpus WikiDBs shows that databases are interconnected, unaligned, and sometimes unjoinable, exposing a significant gap between CL algorithm design and practical deployment. To close this evaluation gap, we build WikiDBGraph, a large-scale dataset constructed from 100,000 real-world relational databases linked by 17 million weighted edges. Each node (database) and edge (relationship) is annotated with 13 and 12 properties, respectively, capturing a hybrid of instance- and feature-level overlap across databases. Experiments on WikiDBGraph demonstrate both the effectiveness and limitations of existing CL methods under realistic conditions, highlighting previously overlooked gaps in managing real-world data silos and pointing to concrete directions for practical deployment of collaborative learning systems.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext bb44935b-d4de-4025-b8ab-6dc4be54d1bbBuilds on22
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- TURL: Table Understanding through Representation LearningXiang Deng, Huan Sun, Alyssa Lees, You Wu et al.VLDB 2021 · 2,406 citations
- MPNet: Masked and Permuted Pre-training for Language UnderstandingKaitao Song, Xu Tan, Tao Qin, Jianfeng Lu et al.NeurIPS 2020 · 1,957 citations
- Federated Learning on Non-IID Data Silos: An Experimental StudyQinbin Li, Yiqun Diao, Quan Chen, Bingsheng HeICDE 2022 · 1,110 citations
- Privacy Preserving Vertical Federated Learning for Tree-based ModelsYuncheng Wu, Shaofeng Cai, Xiaokui Xiao, Gang Chen et al.VLDB 2020 · 259 citations
Related papers
- Subgraph Federated Learning with Missing Neighbor GenerationKe Zhang, Carl Yang, Xiaoxiao Li, Lichao Sun et al.NeurIPS 2021 · 320 citations
- AdaFGL: A New Paradigm for Federated Node Classification with Topology HeterogeneityXunkai Li, Zhengyu Wu, Wentao Zhang, Henan Sun et al.ICDE 2024 · 11 citations
- Learning Over Dirty Data Without CleaningJose Picado, John Davis, Arash Termehchy, Ga Young LeeSIGMOD 2020 · 13 citations
- OpenFGL: A Comprehensive Benchmark for Federated Graph LearningXunkai Li, Yinlin Zhu, Boyang Pang, Guochen Yan et al.VLDB 2025 · 12 citations
- Missingness-aware Federated Contrastive Learning on Semantic GraphsShuo Yu, Zhuoyang Han, Guoqing Han, Tao Tang et al.WWW 2026
