Large Language Model-Enhanced Multi-Armed Bandits
Jiahang Sun, Zhiyong Wang, Runhan Yang, Chenjun Xiao, John C. S. Lui, Zhongxiang Dai
Abstract
Large language models (LLMs) have been adopted to solve sequential decision-making tasks such as multi-armed bandits (MAB), in which an LLM is directly instructed to select the arms to pull in every iteration. However, this paradigm of direct arm selection using LLMs has been shown to be suboptimal in many MAB tasks. Therefore, we propose an alternative approach which combines the strengths of classical MAB and LLMs. Specifically, we adopt a classical MAB algorithm as the high-level framework and leverage the strong in-context learning capability of LLMs to perform the sub-task of reward prediction. Firstly, we incorporate the LLM-based reward predictor into the classical Thompson sampling (TS) algorithm and adopt a decaying schedule for the LLM temperature to ensure a transition from exploration to exploitation. Next, we incorporate the LLM-based reward predictor (with a temperature of 0) into a regression oracle-based MAB algorithm equipped with an explicit exploration mechanism. We also extend our TS-based algorithm to dueling bandits where only the preference feedback between pairs of arms is available, which requires non-trivial algorithmic modifications. We conduct empirical evaluations using both synthetic MAB tasks and experiments designed using real-world text datasets, in which the results show that our algorithms consistently outperform previous baseline methods based on direct arm selection. Interestingly, we also demonstrate that in challenging tasks where the arms lack semantic meanings that can be exploited by the LLM, our approach achieves considerably better performance than LLM-based direct arm selection.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext c43eec61-8fb9-48cd-8ff1-4ff33de96218Cited by top-tier papers3
- Why Keep Your Doubts to Yourself? Trading Visual Uncertainties among Vision-Language ModelsJusheng Zhang, Yijia Fan, Kaitong Cai, Jing Yang et al.ICLR 2026 · 6 citations
- Post-Training LLMs as Better Decision-Making Agents: A Regret-Minimization ApproachChanwoo Park, Ziyang Chen, Asuman Ozdaglar, Kaiqing ZhangICML 2026 · 5 citations
- In-Context Learning for Pure ExplorationAlessio Russo, Ryan Welch, Aldo PacchianoICLR 2026 · 5 citations
Builds on13
- Tree of Thoughts: Deliberate Problem Solving with Large Language ModelsShunyu Yao, Dian Yu, Jeffrey Zhao, Izhak Shafran et al.NeurIPS 2023 · 5,068 citations
- Large Language Models as OptimizersChengrun Yang, Xuezhi Wang, Yifeng Lu, Hanxiao Liu et al.ICLR 2024 · 817 citations
- Beyond UCB: Optimal and Efficient Contextual Bandits with Regression OraclesDylan J. Foster, Alexander RakhlinICML 2020 · 241 citations
- Supervised Pretraining Can Learn In-Context Reinforcement LearningJonathan Lee, Annie Xie, Aldo Pacchiano, Yash Chandak et al.NeurIPS 2023 · 170 citations
- Large Language Models to Enhance Bayesian OptimizationTennison Liu, Nicolás Astorga, Nabeel Seedat, Mihaela van der SchaarICLR 2024 · 143 citations
Related papers
- EVOLvE: Evaluating and Optimizing LLMs For In-Context ExplorationAllen Nie, Yi Su, Bo Chang, Jonathan Lee et al.ICML 2025
- Jump Starting Bandits with LLM-Generated Prior KnowledgeParand A. Alamdari, Yanshuai Cao, Kevin H. WilsonEMNLP 2024 · 1 citation
- Theory of Mind for Multi-Agent Collaboration via Large Language ModelsHuao Li, Yu Quan Chong, Simon Stepputtis, Joseph Campbell et al.EMNLP 2023 · 57 citations
- Efficient Reinforcement Learning with Large Language Model PriorsXue Yan, Yan Song, Xidong Feng, Mengyue Yang et al.ICLR 2025
- Efficient Sequential Decision Making with Large Language ModelsDingyang Chen, Qi Zhang, Yinglun ZhuEMNLP 2024 · 3 citations
