Unlocking the Future: Exploring Look-Ahead Planning Mechanistic Interpretability in Large Language Models
Tianyi Men, Pengfei Cao, Zhuoran Jin, Yubo Chen, Kang Liu, Jun Zhao
摘要
Planning, as the core module of agents, is crucial in various fields such as embodied agents, web navigation, and tool using. With the development of large language models (LLMs), some researchers treat large language models as intelligent agents to stimulate and evaluate their planning capabilities. However, the planning mechanism is still unclear. In this work, we focus on exploring the look-ahead planning mechanism in large language models from the perspectives of information flow and internal representations. First, we study how planning is done internally by analyzing the multi-layer perception (MLP) and multi-head self-attention (MHSA) components at the last token. We find that the output of MHSA in the middle layers at the last token can directly decode the decision to some extent. Based on this discovery, we further trace the source of MHSA by information flow, and we reveal that MHSA mainly extracts information from spans of the goal states and recent steps. According to information flow, we continue to study what information is encoded within it. Specifically, we explore whether future decisions have been encoded in advance in the representation of flow. We demonstrate that the middle and upper layers encode a few shortterm future decisions to some extent when planning is successful. Overall, our research analyzes the look-ahead planning mechanisms of LLMs, facilitating future research on LLMs performing planning tasks.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper14
- On Reasoning Strength Planning in Large Reasoning ModelsLeheng Sheng, An Zhang, Zijian Wu, Weixiang Zhao 等NeurIPS 2025 · 被引用 17 次
- Agent-RewardBench: Towards a Unified Benchmark for Reward Modeling across Perception, Planning, and Safety in Real-World Multimodal AgentsTianyi Men, Zhuoran Jin, Pengfei Cao, Yubo Chen 等ACL 2025 · 被引用 13 次
- Deep sequence models tend to memorize geometrically; it is unclear whyShahriar Noroozizadeh, Vaishnavh Nagarajan, Elan Rosenfeld, Sanjiv KumarICML 2026 · 被引用 11 次
- Language Models Can Predict Their Own BehaviorDhananjay Ashok, Jonathan MayNeurIPS 2025 · 被引用 10 次
- What's the plan? Metrics for implicit planning in LLMs and their application to rhyme generation and question answeringJim Maar, Denis Paperno, Callum McDougall, Neel NandaICLR 2026 · 被引用 6 次
它引用的顶会 Paper15
- Tree of Thoughts: Deliberate Problem Solving with Large Language ModelsShunyu Yao, Dian Yu, Jeffrey Zhao, Izhak Shafran 等NeurIPS 2023 · 被引用 5,068 次
- Language Models as Zero-Shot Planners: Extracting Actionable Knowledge for Embodied AgentsWenlong Huang, Pieter Abbeel, Deepak Pathak, Igor MordatchICML 2022 · 被引用 1,539 次
- ToolLLM: Facilitating Large Language Models to Master 16000+ Real-world APIsYujia Qin, Shihao Liang, Yining Ye, Kunlun Zhu 等ICLR 2024 · 被引用 1,469 次
- WebArena: A Realistic Web Environment for Building Autonomous AgentsShuyan Zhou, Frank F. Xu, Hao Zhu, Xuhui Zhou 等ICLR 2024 · 被引用 1,197 次
- ALFWorld: Aligning Text and Embodied Environments for Interactive LearningMohit Shridhar, Xingdi Yuan, Marc-Alexandre Côté, Yonatan Bisk 等ICLR 2021 · 被引用 819 次
相关 Paper
- Internal Planning in Language Models: Characterizing Horizon and Branch AwarenessMuhammed Ustaomeroglu, Baris Askin, Gauri Joshi, Carlee Joe-Wong 等ICLR 2026
- Interpreting Context Look-ups in Transformers: Investigating Attention-MLP InteractionsClement Neo, Shay B. Cohen, Fazl BarezEMNLP 2024 · 被引用 3 次
- Emergent Response Planning in LLMsZhichen Dong, Zhanhui Zhou, Zhixuan Liu, Chao Yang 等ICML 2025
- Why Do LLM-based Web Agents Fail? A Hierarchical Planning PerspectiveMohamed Aghzal, Gregory J. Stein, Ziyu YaoACL 2026 · 被引用 5 次
- ACTIVE-o3 : Empowering MLLMs with Active Perception via Pure Reinforcement LearningMuzhi Zhu, Hao Zhong, Canyu Zhao, Zongze Du 等ICML 2026 · 被引用 35 次
