AttnPO: Attention-Guided Process Supervision for Efficient Reasoning
Shuaiyi Nie, Siyu Ding, Wenyuan Zhang, Linhao Yu, Tianmeng Yang, Yao Chen, Weichong Yin, Yu Sun, Hua Wu, Tingwen Liu
Abstract
Large reasoning models trained with reinforcement learning and verifiable rewards (RLVR) achieve strong performance on complex reasoning tasks, yet often overthink, generating redundant reasoning without performance gains. Existing trajectory-level length penalties often fail to effectively shorten reasoning length and degrade accuracy, as they uniformly treat all reasoning steps and lack fine-grained signals to distinguish redundancy from necessity. Meanwhile, process-supervised methods are typically resource-intensive and suffer from inaccurate credit assignment. To address these issues, we propose ATTNPO, a low-overhead process-supervised RL framework that leverages the model's intrinsic attention signals for step-level credit assignment. We first identify a set of special attention heads that naturally focus on essential steps while suppressing redundant ones. By leveraging the attention scores of these heads, We then employ two sub-strategies to mitigate overthinking by discouraging redundant steps while preserving accuracy by reducing penalties on essential steps. Experimental results show that ATTNPO substantially reduces reasoning length while significantly improving performance across 9 benchmarks 1 .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 1774d49b-7a7b-42e4-a47b-a65faf399fd2Cited by top-tier papers8
- Scaling Behaviors of LLM Reinforcement Learning Post-Training: An Empirical Study in Mathematical ReasoningZelin Tan, Hejia Geng, Xiaohang Yu, Mulei Zhang et al.ACL 2026 · 17 citations
- Heterogeneous Adaptive Policy Optimization: Tailoring Optimization to Every Token's NatureZheng Liu, Mengjie Liu, Siwei Wen, Mengzhang Cai et al.ACL 2026 · 9 citations
- Reinforced Efficient Reasoning via Semantically Diverse ExplorationZiqi Zhao, Zhaochun Ren, Jiahong Zou, Liu Yang et al.ACL 2026 · 5 citations
- From Curated Data to Scalable Models: Continual Pre-training of Dense and MoE Large Language Models for TibetanLei Yang, Leiyu Pan, Bojian Xiong, Renren Jin et al.ACL 2026 · 5 citations
- SHAPE: Stage-aware Hierarchical Advantage via Potential Estimation for LLM ReasoningZhengyang Ai, Zikang Shan, Xiaodong Ai, Jingxian Tang et al.ACL 2026 · 3 citations
Builds on16
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou et al.ICLR 2021 · 7,905 citations
- Let's Verify Step by StepHunter Lightman, Vineet Kosaraju, Yuri Burda, Harrison Edwards et al.ICLR 2024 · 3,045 citations
- ShorterBetter: Guiding Reasoning Models to Find Optimal Inference Length for Efficient ReasoningJingyang Yi, Jiazheng Wang, Sida LiNeurIPS 2025 · 89 citations
- OpenRubrics: Towards Scalable Synthetic Rubric Generation for Reward Modeling and LLM AlignmentTianci Liu, Ran Xu, Tony Yu, Ilgee Hong et al.ACL 2026 · 75 citations
- HybridFlow: A Flexible and Efficient RLHF FrameworkGuangming Sheng, Chi Zhang, Zilingfeng Ye, Xibin Wu et al.EuroSys 2025 · 61 citations
Related papers
- Incentivizing Dual Process Thinking for Efficient Large Language Model ReasoningXiaoxue Cheng, Junyi Li, Zhenduo Zhang, Xinyu Tang et al.NeurIPS 2025 · 25 citations
- Promoting Efficient Reasoning with Verifiable Stepwise RewardChuhuai Yue, Chengqi Dong, Yinan Gao, Hang He et al.AAAI 2026 · 19 citations
- DRPO: Efficient Reasoning via Decoupled Reward Policy OptimizationGang Li, Yan Chen, Ming Lin, Tianbao YangICLR 2026 · 19 citations
- Overthinking Reduction with Decoupled Rewards and Curriculum Data SchedulingShuyang Jiang, Yusheng Liao, Ya Zhang, Yanfeng Wang et al.ICLR 2026 · 5 citations
- Attention as a Compass: Efficient Exploration for Process-Supervised RL in Reasoning ModelsRunze Liu, Jiakang Wang, Yuling Shi, Zhihui Xie et al.ICLR 2026 · 13 citations
