Empowering LLM Agents with Zero-Shot Optimal Decision-Making through Q-learning
Jiajun Chai, Sicheng Li, Yuqian Fu, Dongbin Zhao, Yuanheng Zhu
摘要
Current Large language model (LLM) agents succeed in making zero-shot decisions but struggle to make optimal decisions, as they rely on pre-trained probabilities rather than maximizing expected future rewards. In contrast, agents trained via reinforcement learning (RL) could make optimal decisions but require extensive data. We develop an algorithm that combines the zero-shot capabilities of LLMs with the optimization of RL, referred to as the Model-based LLM Agent with Q-Learning (MLAQ). MLAQ employs Q-learning to derive optimal policies from transitions within memory. Unlike RL agents, MLAQ constructs an LLM-based imagination space, where a UCB variant generates imaginary data through interactions with the LLM-based world model to derive zero-shot policies. This approach achieves a sub-linear regret bound, as guaranteed by our theorem. Moreover, MLAQ employs a mixed-examination mechanism to further enhance the quality of imaginary data. We evaluate MLAQ on benchmarks that present significant challenges for existing LLM agents. Results show that MLAQ achieves a optimal rate of over 90% in tasks where other methods struggle to succeed. Additional experiments are conducted to reach the conclusion that introducing model-based RL into LLM agents shows significant potential in optimal decision-making. Our website is available at link.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- HeurekaBench: A Benchmarking Framework for AI Co-scientistSiba Smarak Panigrahi, Jovana Videnovic, Maria BrbicICLR 2026 · 被引用 10 次
- RLAE: Reinforcement Learning-Assisted Ensemble for LLMsYuqian Fu, Yuanheng Zhu, Jiajun Chai, Guojun Yin 等EMNLP 2025 · 被引用 1 次
- DipLLM: Fine-Tuning LLM for Strategic Decision-making in DiplomacyKaixuan Xu, Jiajun Chai, Sicheng Li, Yuqian Fu 等ICML 2025
它引用的顶会 Paper13
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- Reflexion: language agents with verbal reinforcement learningNoah Shinn, Federico Cassano, Ashwin Gopinath, Karthik Narasimhan 等NeurIPS 2023 · 被引用 5,828 次
- Tree of Thoughts: Deliberate Problem Solving with Large Language ModelsShunyu Yao, Dian Yu, Jeffrey Zhao, Izhak Shafran 等NeurIPS 2023 · 被引用 5,068 次
- Dream to Control: Learning Behaviors by Latent ImaginationDanijar Hafner, Timothy P. Lillicrap, Jimmy Ba, Mohammad NorouziICLR 2020 · 被引用 1,852 次
相关 Paper
- RILA: Reflective and Imaginative Language Agent for Zero-Shot Semantic Audio-Visual NavigationZeyuan Yang, Jiageng Lin, Peihao Chen, Anoop Cherian 等CVPR 2024 · 被引用 5 次
- Exploratory Memory-Augmented LLM Agent via Hybrid On- and Off-Policy OptimizationZeyuan Liu, Jeonghye Kim, Xufang Luo, Dongsheng Li 等ICLR 2026 · 被引用 18 次
- From Words to Actions: Unveiling the Theoretical Underpinnings of LLM-Driven Autonomous SystemsJianliang He, Siyu Chen, Fengzhuo Zhang, Zhuoran YangICML 2024 · 被引用 12 次
- Large Language Models as Generalizable Policies for Embodied TasksAndrew Szot, Max Schwarzer, Harsh Agrawal, Bogdan Mazoure 等ICLR 2024 · 被引用 114 次
- Reason for Future, Act for Now: A Principled Architecture for Autonomous LLM AgentsZhihan Liu, Hao Hu, Shenao Zhang, Hongyi Guo 等ICML 2024 · 被引用 17 次
