Latent Planning Emerges with Scale
Michael Hanna, Emmanuel Ameisen
摘要
LLMs can perform seemingly planning-intensive tasks, like writing coherent stories or functioning code, without explicitly verbalizing a plan; however, the extent to which they implicitly plan is unknown. In this paper, we define latent planning as occurring when LLMs possess internal planning representations that (1) cause the generation of a specific future token or concept, and (2) shape preceding context to license said future token or concept. We study the Qwen-3 family (0.6B-14B) on simple planning tasks, finding that latent planning ability increases with scale. Models that plan possess features that represent a planned-for word like accountant, and cause them to output an rather than a; moreover, even the less-successful Qwen-3 4B-8B have nascent planning mechanisms. On the more complex task of completing rhyming couplets, we find that models often identify a rhyme ahead of time, but even large models seldom plan far ahead. However, we can elicit some planning that increases with scale when steering models towards planned words in prose. In sum, we offer a framework for measuring planning and mechanistic evidence of how models' planning abilities grow with scale. * Completed as part of the Anthropic Fellows Program. 1 Explicit, verbalized planning, as in LLMs' chains of thought, is better studied (Kambhampati et al., 2024) .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper13
- Sparse Autoencoders Find Highly Interpretable Features in Language ModelsRobert Huben, Hoagy Cunningham, Logan Riggs Smith, Aidan Ewart 等ICLR 2024 · 被引用 1,072 次
- Towards Automated Circuit Discovery for Mechanistic InterpretabilityArthur Conmy, Augustine N. Mavor-Parker, Aengus Lynch, Stefan Heimersheim 等NeurIPS 2023 · 被引用 861 次
- How does GPT-2 compute greater-than?: Interpreting mathematical abilities in a pre-trained language modelMichael Hanna, Ollie Liu, Alexandre VariengienNeurIPS 2023 · 被引用 251 次
- Transcoders find interpretable LLM feature circuitsJacob Dunefsky, Philippe Chlenski, Neel NandaNeurIPS 2024 · 被引用 222 次
- Patchscopes: A Unifying Framework for Inspecting Hidden Representations of Language ModelsAsma Ghandeharioun, Avi Caciularu, Adam Pearce, Lucas Dixon 等ICML 2024 · 被引用 197 次
相关 Paper
- Emergent Response Planning in LLMsZhichen Dong, Zhanhui Zhou, Zhixuan Liu, Chao Yang 等ICML 2025
- Unlocking the Future: Exploring Look-Ahead Planning Mechanistic Interpretability in Large Language ModelsTianyi Men, Pengfei Cao, Zhuoran Jin, Yubo Chen 等EMNLP 2024 · 被引用 20 次
- LLMs Can Plan Only If We Tell ThemBilgehan Sel, Ruoxi Jia, Ming JinICLR 2025
- Internal Planning in Language Models: Characterizing Horizon and Branch AwarenessMuhammed Ustaomeroglu, Baris Askin, Gauri Joshi, Carlee Joe-Wong 等ICLR 2026
- How Far Ahead Do LLMs Plan? Uncovering the Latent Horizon in Chain-of-Thought ReasoningLiyan Xu, Mo Yu, Fandong Meng, Jie ZhouICML 2026 · 被引用 1 次
