Online Mixture of Experts: No-Regret Learning for Optimal Collective Decision-Making
Larkin Liu, Jalal Etesami
Abstract
We explore the use of expert-guided bandit learning, which we refer to as online mixture-of-experts (OMoE). In this setting, given a context, a candidate committee of experts must determine how to aggregate their outputs to achieve optimal results in terms of aggregate accuracy. We propose two algorithms to address this problem. The first algorithm combines aggregate voting with UCB-driven successive elimination, efficiently pruning suboptimal exploration actions. The second algorithm employs an online weighted-majority-voting mechanism, leveraging the respective voting power of each expert proportional to their predictive power. We derive theoretical guarantees for the regret properties in the bandit setting under ideal circumstances, and empirical results are provided accordingly. As a modern study on applications, these methods are applied to the online fine-tuning of a set of expert large language models (LLMs), where after each response, the generative LLM dynamically reweighs its set of experts and/or selects the optimal committee of experts to generate the most accurate response. Our results introduce new methodologies and no-regret guarantees for combining multiple experts to improve on the performance of the an aggregate model overall.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on4
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- DeepSeekMoE: Towards Ultimate Expert Specialization in Mixture-of-Experts Language ModelsDamai Dai, Chengqi Deng, Chenggang Zhao, R. X. Xu et al.ACL 2024 · 171 citations
- Multi-Head Mixture-of-ExpertsXun Wu, Shaohan Huang, Wenhui Wang, Shuming Ma et al.NeurIPS 2024 · 42 citations
- Satisficing Regret Minimization in BanditsQing Feng, Tianyi Ma, Ruihao ZhuICLR 2025 · 1 citation
Related papers
- Beyond Majority Voting: LLM Aggregation by Leveraging Higher-Order InformationRui Ai, Yuqi Pan, David Simchi-Levi, Milind Tambe et al.ICML 2026 · 20 citations
- Unveiling and Consulting Core Experts in Retrieval-Augmented MoE-based LLMsXin Zhou, Ping Nie, Yiwen Guo, Haojie Wei et al.EMNLP 2024
- Best-of-Infinity: Asymptotic Performance of Test-Time LLM EnsemblingJunpei Komiyama, Daisuke Oba, Masafumi OyamadaICLR 2026 · 2 citations
- Filtered not Mixed: Filtering-Based Online Gating for Mixture of Large Language ModelsRaeid Saqur, Anastasis Kratsios, Florian Krach, Yannick Limmer et al.ICLR 2025
- Your Mixture-of-Experts LLM Is Secretly an Embedding Model for FreeZiyue Li, Tianyi ZhouICLR 2025
