Bootstrapping Self-Improvement of Language Model Programs for Zero-Shot Schema Matching
Nabeel Seedat, Mihaela van der Schaar
Abstract
Schema matching -the task of finding matches between attributes across disparate data sources with different tables and hierarchies -is critical for creating interoperable machine learning (ML)-ready data. Addressing this fundamental data-centric problem has wide implications, especially in domains like healthcare, finance and e-commercebut also has the potential to benefit ML models more generally, by increasing the data available for ML model training. However, schema matching is a challenging ML task due to structural/hierarchical and semantic heterogeneity between different schemas. Previous ML approaches to automate schema matching have either required significant labeled data for model training, which is often unrealistic or suffer from poor zero-shot performance. To this end, we propose Matchmaker -a compositional language model program for schema matching, comprised of candidate generation, refinement and confidence scoring. Matchmaker also self-improves in a zero-shot manner without the need for labeled demonstrations via a novel optimization approach, which constructs synthetic in-context demonstrations to guide the language model's reasoning process. Empirically, we demonstrate on real-world medical schema matching benchmarks that Matchmaker outperforms previous ML-based approaches, highlighting its potential to accelerate data integration and interoperability of ML-ready data.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 8face420-c55b-498d-8133-28b518f89e0eBuilds on17
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- Self-Refine: Iterative Refinement with Self-FeedbackAman Madaan, Niket Tandon, Prakhar Gupta, Skyler Hallinan et al.NeurIPS 2023 · 4,972 citations
- "Everyone wants to do the model work, not the data work": Data Cascades in High-Stakes AINithya Sambasivan, Shivani Kapania, Hannah Highfill, Diana Akrong et al.CHI 2021 · 725 citations
- Deep Entity Matching with Pre-Trained Language ModelsYuliang Li, Jinfeng Li, Yoshihiko Suhara, AnHai Doan et al.VLDB 2021 · 484 citations
Related papers
- Schema Matching using Pre-Trained Language ModelsYunjia Zhang, Avrilia Floratou, Joyce Cahoon, Subru Krishnan et al.ICDE 2023 · 32 citations
- In Situ Neural Relational Schema MatcherXingyu Du, Gongsheng Yuan, Sai Wu, Gang Chen et al.ICDE 2024 · 4 citations
- CoMCo: Consistency-Aware Multi-Agent Coordination for Zero-Shot Cross-Modal Entity MatchingShiqi Zhang, Weixin Zeng, Ziheng Zhang, Wenzhe Hou et al.SIGIR 2026
- Magneto: Combining Small and Large Language Models for Schema MatchingYurong Liu, Eduardo H. M. Pena, Aécio S. R. Santos, Eden Wu et al.VLDB 2025 · 32 citations
- Agent-OM: Leveraging LLM Agents for Ontology MatchingZhangcheng Qiang, Weiqing Wang, Kerry TaylorVLDB 2025 · 34 citations
