Principled Model Routing for Unknown Mixtures of Source Domains
Christoph Dann, Yishay Mansour, Teodor Vanislavov Marinov, Mehryar Mohri
Abstract
The rapid proliferation of domain-specialized machine learning models presents a challenge: while individual models excel in specific domains, their performance varies significantly across diverse applications. This makes selecting the optimal model when faced with an unknown mixture of tasks, especially with limited or no data to estimate the mixture, a difficult problem. We address this challenge by formulating it as a multiple-source domain adaptation (MSA) problem. We introduce a novel, scalable algorithm that effectively routes each input to the best-suited model from a pool of available models. Our approach provides a strong performance guarantee: remarkably, for any mixture domain, the accuracy achieved by the best source model is maintained. This guarantee is established through a theoretical bound on the regret for new domains, expressed as a convex combination of the best regrets in the source domains, plus a concentration term that diminishes as the amount of source data increases. While our primary contributions are theoretical and algorithmic, we also present empirical results demonstrating the effectiveness of our approach.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 7b8ce0dc-84da-4711-88ba-3be81f44dea8Cited by top-tier papers1
Ask how each one uses itBuilds on14
- Model soups: averaging weights of multiple fine-tuned models improves accuracy without increasing inference timeMitchell Wortsman, Gabriel Ilharco, Samir Yitzhak Gadre, Rebecca Roelofs et al.ICML 2022 · 1,464 citations
- Adversarially Trained Actor Critic for Offline Reinforcement LearningChing-An Cheng, Tengyang Xie, Nan Jiang, Alekh AgarwalICML 2022 · 156 citations
- Large Language Model Cascades with Mixture of Thought Representations for Cost-Efficient ReasoningMurong Yue, Jie Zhao, Min Zhang, Liang Du et al.ICLR 2024 · 153 citations
- AutoMix: Automatically Mixing Language ModelsPranjal Aggarwal, Aman Madaan, Ankit Anand, Srividya Pranavi Potharaju et al.NeurIPS 2024 · 145 citations
- Two-Stage Learning to Defer with Multiple ExpertsAnqi Mao, Christopher Mohri, Mehryar Mohri, Yutao ZhongNeurIPS 2023 · 98 citations
Related papers
- A Discriminative Technique for Multiple-Source AdaptationCorinna Cortes, Mehryar Mohri, Ananda Theertha Suresh, Ningshan ZhangICML 2021 · 15 citations
- Aggregating From Multiple Target-Shifted SourcesChangjian Shui, Zijian Li, Jiaqi Li, Christian Gagné et al.ICML 2021 · 36 citations
- Unsupervised Multi-Source Domain Adaptation Without Access to Source DataSk Miraj Ahmed, Dripta S. Raychaudhuri, Sujoy Paul, Samet Oymak et al.CVPR 2021
- Mixture Weight Estimation and Model Prediction in Multi-source Multi-target Domain AdaptationYuyang Deng, Ilja Kuzborskij, Mehrdad MahdaviNeurIPS 2023 · 3 citations
- Domain Aggregation Networks for Multi-Source Domain AdaptationJunfeng Wen, Russell Greiner, Dale SchuurmansICML 2020 · 82 citations
