Attention Illuminates LLM Reasoning: The Uncovered Preplan-and-Anchor Rhythm Enables Fine-Grained Policy Optimization
Yang Li, Zhichen Dong, Yuhan Sun, Weixun Wang, Shaopan Xiong, Yijia Luo, Jiashun Liu, Han Lu, Jiamang Wang, Wenbo Su, Bo Zheng, Junchi Yan
Abstract
The reasoning patterns of large language models (LLMs) remain opaque, and reinforcement learning (RL) typically assigns uniform credit across an entire generation, blurring the distinction between pivotal and routine steps. This work treats attention as a natural substrate for interpreting LLM reasoning and a window for aligning optimization with its internal dynamics. We first distinguish attention heads between locally and globally focused information processing and reveal that locally focused heads produce a sawtooth pattern near the diagonal indicating phrasal chunks, while globally focused heads expose tokens that exert broad downstream influence over future tokens. We quantify these with two metrics measuring the extent of backward attention within a clipped window and the average attention a token receives from subsequent tokens, respectively. Taken together, these signals indicate a recurring preplan-and-anchor regularity, where the model first performs a long-range contextual reference to generate an introductory token, which is immediately followed by or coincides with a semantic anchor token that organizes subsequent reasoning. Leveraging these insights, we introduce three novel RL strategies that dynamically perform targeted credit assignment to critical nodes (preplan tokens, anchor tokens, and their temporal coupling) and show consistent performance gains across various reasoning tasks. By aligning optimization with the model's intrinsic reasoning rhythm, we aim to transform opaque optimization into an actionable structure-aware process, hoping to offer a potential step toward more transparent and effective optimization of LLM reasoning. Focus More on Critical Nodes of the Reasoning Path This model finally finds more hidden patterns 𝐷 !"#$ = 1 𝑁 % % 𝑨 %,' ' (𝑖 -𝑗) %() '*)
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers4
- Generation as Search Operator for Test-Time Scaling of Diffusion-based Combinatorial OptimizationYang Li, Lvda Chen, Haonan Wang, Runzhong Wang et al.NeurIPS 2025 · 13 citations
- Retro-R1: LLM-based Agentic RetrosynthesisWei Liu, Jiangtao Feng, Hongli Yu, Yuxuan Song et al.NeurIPS 2025 · 8 citations
- Reasoning Palette: Modulating Reasoning via Latent Contextualization for Controllable Exploration for (V)LMsRujiao Long, Yang Li, Xingyao Zhang, Weixun Wang et al.CVPR 2026 · 2 citations
- Beyond Logits: Metastable Latent Dynamics for Sample-Efficient Best-of-N Selection in LLMsXinrong Li, Zidong Zhou, Keyu Shen, Wenhao Zhou et al.ICML 2026
Builds on31
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning et al.NeurIPS 2023 · 10,924 citations
- FlashAttention: Fast and Memory-Efficient Exact Attention with IO-AwarenessTri Dao, Daniel Y. Fu, Stefano Ermon, Atri Rudra et al.NeurIPS 2022 · 5,493 citations
- DAPO: An Open-Source LLM Reinforcement Learning System at ScaleQiying Yu, Zheng Zhang, Ruofei Zhu, Yufeng Yuan et al.NeurIPS 2025 · 2,828 citations
- Efficient Streaming Language Models with Attention SinksGuangxuan Xiao, Yuandong Tian, Beidi Chen, Song Han et al.ICLR 2024 · 1,714 citations
Related papers
- How Does Reasoning Flow? Tracing Attention-Induced Information Flow for Targeted RL in LLMsZhichen Dong, Yang Li, Yuhan Sun, Weixun Wang et al.ICML 2026 · 1 citation
- Emergent Hierarchical Reasoning in LLMs through Reinforcement LearningHaozhe Wang, Qixin Xu, Che Liu, Junhong Wu et al.ICLR 2026 · 44 citations
- GTPO and GRPO-S: Token and Sequence-Level Reward Shaping with Policy EntropyHongze Tan, Zihan Wang, Jianfei Pan, Jinghao Lin et al.ICML 2026 · 53 citations
- Reshaping Reasoning in LLMs: A Theoretical Analysis of RL Training Dynamics through Pattern SelectionXingwu Chen, Tianle Li, Difan ZouICLR 2026 · 8 citations
- Beyond Fully Random Masking: Attention-Guided Denoising and Optimization for Diffusion Language ModelsJia Deng, Junyi Li, Xin Zhao, Jinpeng Wang et al.ACL 2026
