Online Multi-LLM Selection via Contextual Bandits Under Unstructured Context Evolution
Manhin Poon, Xiangxiang Dai, Xutong Liu, Fang Kong, John C. S. Lui, Jinhang Zuo
Abstract
Large language models (LLMs) exhibit diverse response behaviors, costs, and strengths, making it challenging to select the most suitable LLM for a given user query. We study the problem of adaptive multi-LLM selection in an online setting, where the learner interacts with users through multi-step query refinement and must choose LLMs sequentially without access to offline datasets or model internals. A key challenge arises from unstructured context evolution: the prompt dynamically changes in response to previous model outputs via a black-box process, which cannot be simulated, modeled, or learned. To address this, we propose the first contextual bandit framework for sequential LLM selection under unstructured prompt dynamics. We formalize a notion of myopic regret and develop a LinUCB-based algorithm that provably achieves sublinear regret without relying on future context prediction. We further introduce budget-aware and positionally-aware (favoring early-stage satisfaction) extensions to accommodate variable query costs and user preferences for early high-quality responses. Our algorithms are theoretically grounded and require no offline fine-tuning or dataset-specific training. Experiments on diverse benchmarks demonstrate that our methods outperform existing LLM routing strategies in both accuracy and cost-efficiency, validating the power of contextual bandits for real-time, adaptive LLM selection. Code is available at https://github.com/EntroShape/Online_LLM_Selection
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 89f01246-71fb-4839-85ad-836758ea4a89Cited by top-tier papers3
- SLM-MUX: Orchestrating Small Language Models for ReasoningChenyu Wang, Zishen Wan, Hao Kang, Emma Chen et al.ICLR 2026 · 9 citations
- Near-Optimal Online Deployment and Routing for Streaming LLMsShaoang Li, Jian LiICLR 2026 · 3 citations
- Ghost in the Transformer: Detecting Model Reuse with Invariant Spectral SignaturesSuqing Wang, Ziyang Ma, Xinyi Li, Zuchao LiAAAI 2026 · 1 citation
Builds on7
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Tree of Thoughts: Deliberate Problem Solving with Large Language ModelsShunyu Yao, Dian Yu, Jeffrey Zhao, Izhak Shafran et al.NeurIPS 2023 · 5,068 citations
- Self-Rewarding Language ModelsWeizhe Yuan, Richard Yuanzhe Pang, Kyunghyun Cho, Xian Li et al.ICML 2024 · 569 citations
- PromptAgent: Strategic Planning with Language Models Enables Expert-level Prompt OptimizationXinyuan Wang, Chenxi Li, Zhen Wang, Fan Bai et al.ICLR 2024 · 226 citations
- AutoMix: Automatically Mixing Language ModelsPranjal Aggarwal, Aman Madaan, Ankit Anand, Srividya Pranavi Potharaju et al.NeurIPS 2024 · 145 citations
Related papers
- BARouter: A Budget-adaptive Online Large Language Model Router FrameworkLingkai Zu, Xiyue Peng, Xin LiuWWW 2026
- Jump Starting Bandits with LLM-Generated Prior KnowledgeParand A. Alamdari, Yanshuai Cao, Kevin H. WilsonEMNLP 2024 · 1 citation
- Efficient Multi-objective Prompt Optimization via Pure-exploration BanditsDonghao Li, Chengshuai Shi, Weijuan Ou, Cong Shen et al.ICLR 2026 · 2 citations
- Efficient Prompt Optimization Through the Lens of Best Arm IdentificationChengshuai Shi, Kun Yang, Zihan Chen, Jundong Li et al.NeurIPS 2024 · 44 citations
- CES: Combinatorial Experts Selection via Contextual Linear BanditsJinkun Xu, Minghan Wang, Zhiyong Wang, Zhongxiang Dai et al.KDD 2026
