From Words to Actions: Unveiling the Theoretical Underpinnings of LLM-Driven Autonomous Systems
Jianliang He, Siyu Chen, Fengzhuo Zhang, Zhuoran Yang
摘要
In this work, from a theoretical lens, we aim to understand why large language model (LLM) empowered agents are able to solve decision-making problems in the physical world. To this end, consider a hierarchical reinforcement learning (RL) model where the LLM Planner and the Actor perform high-level task planning and low-level execution, respectively. Under this model, the LLM Planner navigates a partially observable Markov decision process (POMDP) by iteratively generating language-based subgoals via prompting. Under proper assumptions on the pretraining data, we prove that the pretrained LLM Planner effectively performs Bayesian aggregated imitation learning (BAIL) through in-context learning. Additionally, we highlight the necessity for exploration beyond the subgoals derived from BAIL by proving that naively executing the subgoals returned by LLM leads to a linear regret. As a remedy, we introduce an -greedy exploration strategy to BAIL, which is proven to incur sublinear regret when the pretraining error is small. Finally, we extend our theoretical framework to include scenarios where the LLM Planner serves as a world model for inferring the transition model of the environment and to multi-agent settings, enabling coordination among multiple Actors.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- Tru-POMDP: Task Planning Under Uncertainty via Tree of Hypotheses and Open-Ended POMDPsWenjing Tang, Xinyu He, Yongxi Huang, Yunxiao Xiao 等NeurIPS 2025 · 被引用 6 次
- Principle-Evolvable Scientific Discovery via Uncertainty MinimizationYingming Pu, Tao LIN, Hongyu ChenICML 2026 · 被引用 4 次
- SkillGen: Learning Domain Skills for In-Context Sequential Decision MakingRuomeng Ding, Wei Cheng, Minglai Shao, Chen ZhaoAAAI 2026
- Assessing Safety Risks and Quantization-aware Safety Patching for Quantized Large Language ModelsKejia Chen, Jiawen Zhang, Jiacong Hu, Yu Wang 等ICML 2025
- In-Context Linear Regression Demystified: Training Dynamics and Mechanistic Interpretability of Multi-Head Softmax AttentionJianliang He, Xintian Pan, Siyu Chen, Zhuoran YangICML 2025
它引用的顶会 Paper28
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- Scaling Up Visual and Vision-Language Representation Learning With Noisy Text SupervisionChao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen 等ICML 2021 · 被引用 5,401 次
- Tree of Thoughts: Deliberate Problem Solving with Large Language ModelsShunyu Yao, Dian Yu, Jeffrey Zhao, Izhak Shafran 等NeurIPS 2023 · 被引用 5,068 次
相关 Paper
- Reason for Future, Act for Now: A Principled Architecture for Autonomous LLM AgentsZhihan Liu, Hao Hu, Shenao Zhang, Hongyi Guo 等ICML 2024 · 被引用 17 次
- Model-Based Imaginative Planning for Embodied AgentsJunru Song, Hengzhe Jin, Yucong Huang, Tingsong Jiang 等ACL 2026
- Strategic Planning: A Top-Down Approach to Option GenerationMax Ruiz Luyten, Antonin Berthon, Mihaela van der SchaarICML 2025
- Why Do LLM-based Web Agents Fail? A Hierarchical Planning PerspectiveMohamed Aghzal, Gregory J. Stein, Ziyu YaoACL 2026 · 被引用 5 次
- Efficient Reinforcement Learning with Large Language Model PriorsXue Yan, Yan Song, Xidong Feng, Mengyue Yang 等ICLR 2025
