FedAugment: Table Augmentation Search over Decentralized Data Repositories
Lennart Behme, Emil Badura, Leonard Geißler, Matthias Böhm, Ziawasch Abedjan, Volker Markl
Abstract
Dataset search often aims to identify joinable or unionable datasets to augment a given query table. State-of-the-art approaches rely on large language models (LLMs) to embed tables into vector representations and perform semantic similarity search. However, existing work assumes a centralized data repository with embeddings generated by a single, homogeneous pipeline. In contrast to this simplifying assumption, data repositories in the real world are decentralized across multiple data providers, each operating their own embedding pipelines. Given the rapid pace of LLM development and provider-specific fine-tuning, enforcing a standardized pipeline is unrealistic. We introduce FedAugment, a framework for table augmentation search over decentralized data repositories with heterogeneous embeddings. FedAugment constructs a representative set of training examples, embeds it using the individual providers' pipelines, and learns projection functions that align heterogeneous embeddings into a shared vector space via multi-view contrastive learning. Using these projections, all embeddings are mapped into a globally aligned space that supports unified vector similarity search. Compared to issuing independent top- k queries to each data provider, FedAugment enables the retrieval of a global top- k result across all repositories, avoiding redundant retrievals and enabling cost-efficient table augmentation search in decentralized settings.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b737e6b7-2c41-4a31-881b-2754f0fb7d20Builds on32
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- TURL: Table Understanding through Representation LearningXiang Deng, Huan Sun, Alyssa Lees, You Wu et al.VLDB 2021 · 2,406 citations
- TaBERT: Pretraining for Joint Understanding of Textual and Tabular DataPengcheng Yin, Graham Neubig, Wen-tau Yih, Sebastian RiedelACL 2020 · 417 citations
- Matryoshka Representation LearningAditya Kusupati, Gantavya Bhatt, Aniket Rege, Matthew Wallingford et al.NeurIPS 2022 · 364 citations
Related papers
- Featurized-Decomposition Join: Low-Cost Semantic Joins with GuaranteesSepanta Zeighami, Shreya Shankar, Aditya G. ParameswaranVLDB 2026 · 11 citations
- Exploring Representation-level Augmentation for Code SearchHaochen Li, Chunyan Miao, Cyril Leung, Yanxian Huang et al.EMNLP 2022 · 17 citations
- Training Dense Retrievers with Multiple Positive PassagesBenben Wang, Minghao Tang, Hengran Zhang, Jiafeng Guo et al.KDD 2026 · 1 citation
- Multi-Facet Blending for Faceted Query-by-Example RetrievalHeejin Do, Sangwon Ryu, Jonghwi Kim, Gary LeeACL 2025
- Enhancing Visual Representation with Textual Semantics: Textual Semantics-Powered Prototypes for Heterogeneous Federated LearningXinghao Wu, Jianwei Niu, Xuefeng Liu, Guogang Zhu et al.CVPR 2026 · 4 citations
