CES: Combinatorial Experts Selection via Contextual Linear Bandits
Jinkun Xu, Minghan Wang, Zhiyong Wang, Zhongxiang Dai, Fang Kong
Abstract
With the rapid advancement of large language models (LLMs), multi-agent systems have emerged as a promising alternative to scaling up a single model. Existing approaches ensemble multiple LLMs to improve response quality, but they often rely on static prior knowledge of model capabilities and prompts, and require extensive parameter tuning. Some of the methods also treat each combination of LLMs as a learning objective, which leads to exponential time complexity. In this work, we propose an offline-to-online combinatorial experts selection (CES) framework to address these limitations. CES leverages offline evaluation to warm-start model capability estimation and employs online learning to adapt to capability shifts and correct offline inaccuracies. By integrating model features and input semantic representations into a combinatorial multi-armed bandit formulation, CES captures the interaction between prompts and LLMs without introducing complex auxiliary structures such as knowledge graphs. Modeling each LLM as a base arm with answer quality represented by a linear function of model and prompt features, CES achieves polynomial time complexity while excellently balancing performance and cost. Our experiments, conducted on popular LLM evaluation datasets such as AlpacaEval 2.0, show CES's effectiveness, laying the groundwork for future extensions.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get f221c4a7-116c-4029-a95e-e2e9efc61189Related papers
- Online Multi-LLM Selection via Contextual Bandits Under Unstructured Context EvolutionManhin Poon, Xiangxiang Dai, Xutong Liu, Fang Kong et al.AAAI 2026 · 11 citations
- Large Language Model-Enhanced Multi-Armed BanditsJiahang Sun, Zhiyong Wang, Runhan Yang, Chenjun Xiao et al.ACL 2026 · 6 citations
- Mixture-of-Agents Enhances Large Language Model CapabilitiesJunlin Wang, Jue Wang, Ben Athiwaratkun, Ce Zhang et al.ICLR 2025
- Efficient Sequential Decision Making with Large Language ModelsDingyang Chen, Qi Zhang, Yinglun ZhuEMNLP 2024 · 3 citations
- Cost-efficient Knowledge-based Question Answering with Large Language ModelsJunnan Dong, Qinggang Zhang, Chuang Zhou, Hao Chen et al.NeurIPS 2024 · 10 citations
