Can large language models explore in-context?
Akshay Krishnamurthy, Keegan Harris, Dylan J. Foster, Cyril Zhang, Aleksandrs Slivkins
Abstract
We investigate the extent to which contemporary Large Language Models (LLMs) can engage in exploration, a core capability in reinforcement learning and decision making. We focus on native performance of existing LLMs, without training interventions. We deploy LLMs as agents in simple multi-armed bandit environments, specifying the environment description and interaction history entirely in-context, i.e., within the LLM prompt. We experiment with GPT-3.5, GPT-4, and Llama2, using a variety of prompt designs, and find that the models do not robustly engage in exploration without substantial interventions: i) Across all of our experiments, only one configuration resulted in satisfactory exploratory behavior: GPT-4 with chain-of-thought reasoning and an externally summarized interaction history, presented as sufficient statistics; ii) All other configurations did not result in robust exploratory behavior, including those with chain-of-thought reasoning but unsummarized history. Although these findings can be interpreted positively, they suggest that external summarization -- which may not be possible in more complex settings -- is important for obtaining desirable behavior from LLM agents. We conclude that non-trivial algorithmic interventions, such as fine-tuning or dataset curation, may be required to empower LLM-based decision making agents in complex settings.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 74aec1e2-f3aa-4b2a-8d81-50092dd4fa27Cited by top-tier papers26
- Reward Is Enough: LLMs Are In-Context Reinforcement LearnersKefan Song, Amir Moeini, Peng Wang, Lei Gong et al.ICLR 2026 · 42 citations
- AutoDiscovery: Open-ended Scientific Discovery via Bayesian SurpriseDhruv Agarwal, Bodhisattwa Prasad Majumder, Reece Adamson, Megha Chakravorty et al.NeurIPS 2025 · 35 citations
- Meta-RL Induces Exploration in Language AgentsYulun Jiang, Liangze Jiang, Damien Teney, Michael Moor et al.ICLR 2026 · 20 citations
- Toward Efficient Exploration by Large Language Model AgentsDilip Arumugam, Thomas L. GriffithsICLR 2026 · 17 citations
- Large Language Models Think Too Fast To Explore EffectivelyLan Pan, Hanbo Xie, Robert C. WilsonNeurIPS 2025 · 10 citations
Builds on36
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- Large Language Models are Zero-Shot ReasonersTakeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo et al.NeurIPS 2022 · 8,168 citations
- Reflexion: language agents with verbal reinforcement learningNoah Shinn, Federico Cassano, Ashwin Gopinath, Karthik Narasimhan et al.NeurIPS 2023 · 5,828 citations
- Generative Agents: Interactive Simulacra of Human BehaviorJoon Sung Park, Joseph C. O'Brien, Carrie Jun Cai, Meredith Ringel Morris et al.UIST 2023 · 1,882 citations
Related papers
- Why Johnny Can't Prompt: How Non-AI Experts Try (and Fail) to Design LLM PromptsJ. D. Zamfirescu-Pereira, Richmond Y. Wong, Bjoern Hartmann, Qian YangCHI 2023 · 892 citations
- EVOLvE: Evaluating and Optimizing LLMs For In-Context ExplorationAllen Nie, Yi Su, Bo Chang, Jonathan Lee et al.ICML 2025
- AgentBench: Evaluating LLMs as AgentsXiao Liu, Hao Yu, Hanchen Zhang, Yifan Xu et al.ICLR 2024 · 748 citations
- AgentRefine: Enhancing Agent Generalization through Refinement TuningDayuan Fu, Keqing He, Yejie Wang, Wentao Hong et al.ICLR 2025
- Efficient Multi-objective Prompt Optimization via Pure-exploration BanditsDonghao Li, Chengshuai Shi, Weijuan Ou, Cong Shen et al.ICLR 2026 · 2 citations
