Fusing Models with Complementary Expertise
Hongyi Wang, Felipe Maia Polo, Yuekai Sun, Souvik Kundu, Eric P. Xing, Mikhail Yurochkin
摘要
Training AI models that generalize across tasks and domains has long been among the open problems driving AI research. The emergence of Foundation Models made it easier to obtain expert models for a given task, but the heterogeneity of data that may be encountered at test time often means that any single expert is insufficient. We consider the Fusion of Experts (FoE) problem of fusing outputs of expert models with complementary knowledge of the data distribution and formulate it as an instance of supervised learning. Our method is applicable to both discriminative and generative tasks and leads to significant performance improvements in image and text classification, text summarization, multiple-choice QA, and automatic evaluation of generated text. We also extend our method to the "frugal" setting where it is desired to reduce the number of expert model evaluations at test time. Our implementation is publicly available at https: //github.com/hwang595/FoE-ICLR2024 .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper21
- Twin-Merging: Dynamic Integration of Modular Expertise in Model MergingZhenyi Lu, Chenghao Fan, Wei Wei, Xiaoye Qu 等NeurIPS 2024 · 被引用 139 次
- RouterDC: Query-Based Router by Dual Contrastive Learning for Assembling Large Language ModelsShuhao Chen, Weisen Jiang, Baijiong Lin, James T. Kwok 等NeurIPS 2024 · 被引用 113 次
- Ensemble Learning for Heterogeneous Large Language Models with Deep Parallel CollaborationYichong Huang, Xiaocheng Feng, Baohang Li, Yang Xiang 等NeurIPS 2024 · 被引用 94 次
- Smoothie: Label Free Language Model RoutingNeel Guha, Mayee F. Chen, Trevor Chow, Ishan S. Khare 等NeurIPS 2024 · 被引用 44 次
- When One LLM Drools, Multi-LLM Collaboration RulesShangbin Feng, Wenxuan Ding, Alisa Liu, Zifeng Wang 等ACL 2026 · 被引用 27 次
它引用的顶会 Paper19
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou 等ICLR 2021 · 被引用 7,905 次
- PEGASUS: Pre-training with Extracted Gap-sentences for Abstractive SummarizationJingqing Zhang, Yao Zhao, Mohammad Saleh, Peter J. LiuICML 2020 · 被引用 2,453 次
- WILDS: A Benchmark of in-the-Wild Distribution ShiftsPang Wei Koh, Shiori Sagawa, Henrik Marklund, Sang Michael Xie 等ICML 2021 · 被引用 1,773 次
- Ensemble Distillation for Robust Model Fusion in Federated LearningTao Lin, Lingjing Kong, Sebastian U. Stich, Martin JaggiNeurIPS 2020 · 被引用 1,615 次
相关 Paper
- Learning Task-Agnostic Embedding of Multiple Black-Box Experts for Multi-Task Model FusionTrong Nghia Hoang, Thanh Lam, Bryan Kian Hsiang Low, Patrick JailletICML 2020 · 被引用 22 次
- How to Merge Your Multimodal Models Over Time?Sebastian Dziadzio, Vishaal Udandarao, Karsten Roth, Ameya Prabhu 等CVPR 2025
- EmBrace: A Collective Knowledge Fusion Framework Toward Unified EEG Foundation ModelsChenyu Liu, MUYUN JIANG, Pu Wan, Jinxin Pi 等ICML 2026
- A Graph Foundation Model with Cross-Modal Alignment and Modality-Aware Expert Fusion for Multi-Modal GraphsDongxiao He, AnKang Yang, Jitao Zhao, Di JinICML 2026
- FEDKIM: Adaptive Federated Knowledge Injection into Medical Foundation ModelsXiaochen Wang, Jiaqi Wang, Houping Xiao, Jinghui Chen 等EMNLP 2024 · 被引用 6 次
