Magneto: Combining Small and Large Language Models for Schema Matching
Yurong Liu, Eduardo H. M. Pena, Aécio S. R. Santos, Eden Wu, Juliana Freire
摘要
Recent advances in language models (LMs) open new opportunities for schema matching (SM). Recent approaches have shown their potential and key limitations: while small LMs (SLMs) require costly, difficult-to-obtain training data, large LMs (LLMs) demand significant computational resources and face context window constraints. We present Magneto, a cost-effective and accurate solution for SM that combines the advantages of SLMs and LLMs to address their limitations. By structuring the SM pipeline in two phases, retrieval and reranking, Magneto can use computationally efficient SLM-based strategies to derive candidate matches which can then be reranked by LLMs, thus making it possible to reduce runtime while improving matching accuracy. We propose (1) a self-supervised approach to fine-tune SLMs which uses LLMs to generate syntactically diverse training data, and (2) prompting strategies that are effective for reranking. We also introduce a new benchmark, developed in collaboration with domain experts, which includes real biomedical datasets and presents new challenges for SM methods. Through a detailed experimental evaluation, using both our new and existing benchmarks, we show that Magneto is scalable and attains high accuracy for datasets from different domains.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper11
- Unveiling Challenges for LLMs in Enterprise Data EngineeringJan-Micha Bodensohn, Ulf Brackmann, Liane Vogel, Anupam Sanghi 等VLDB 2026 · 被引用 13 次
- QUEST: Query Optimization in Unstructured Document AnalysisZhaoze Sun, Chengliang Chai, Qiyan Deng, Kaisen Jin 等VLDB 2025 · 被引用 9 次
- CENTS: A Flexible and Cost-Effective Framework for LLM-Based Table UnderstandingGuorui Xiao, Dong He, Jin Wang, Magdalena BalazinskaVLDB 2025 · 被引用 4 次
- OmniMatch: Joinability Discovery in Data ProductsChristos Koutras, Jiani Zhang, Xiao Qin, Chuan Lei 等VLDB 2025 · 被引用 3 次
- BDIViz: An Interactive Visualization System for Biomedical Schema Matching with LLM-Powered ValidationEden Wu, Dishita G. Turakhia, Guande Wu, Christos Koutras 等IEEE VIS 2025 · 被引用 3 次
它引用的顶会 Paper17
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- MPNet: Masked and Permuted Pre-training for Language UnderstandingKaitao Song, Xu Tan, Tao Qin, Jianfeng Lu 等NeurIPS 2020 · 被引用 1,957 次
- A Survey on In-context LearningQingxiu Dong, Lei Li, Damai Dai, Ce Zheng 等EMNLP 2024 · 被引用 479 次
- TaBERT: Pretraining for Joint Understanding of Textual and Tabular DataPengcheng Yin, Graham Neubig, Wen-tau Yih, Sebastian RiedelACL 2020 · 被引用 417 次
- Creating Embeddings of Heterogeneous Relational Datasets for Data Integration TasksRiccardo Cappuzzo, Paolo Papotti, Saravanan ThirumuruganathanSIGMOD 2020 · 被引用 139 次
相关 Paper
- UDAPDR: Unsupervised Domain Adaptation via LLM Prompting and Distillation of RerankersJon Saad-Falcon, Omar Khattab, Keshav Santhanam, Radu Florian 等EMNLP 2023 · 被引用 23 次
- Bootstrapping Self-Improvement of Language Model Programs for Zero-Shot Schema MatchingNabeel Seedat, Mihaela van der SchaarICML 2025
- RefineBench: Evaluating Refinement Capability of Language Models via ChecklistsYoung-Jun Lee, Seungone Kim, Byung-Kwan Lee, Minkyeong Moon 等ICLR 2026 · 被引用 13 次
- Schema Matching using Pre-Trained Language ModelsYunjia Zhang, Avrilia Floratou, Joyce Cahoon, Subru Krishnan 等ICDE 2023 · 被引用 32 次
- SciBench: Evaluating College-Level Scientific Problem-Solving Abilities of Large Language ModelsXiaoxuan Wang, Ziniu Hu, Pan Lu, Yanqiao Zhu 等ICML 2024 · 被引用 220 次
