Segment-Level Attribution for Selective Learning of Long Reasoning Traces
Siyuan Wang, Yanchen Liu, Xiang Ren
Abstract
Large Reasoning Models (LRMs) achieve strong reasoning performance by generating long chains of thought (CoTs), yet only a small fraction of these traces meaningfully contributes to answer prediction, while the majority contains repetitive or truncated content. Such output redundancy is further propagated after supervised finetuning (SFT), as models learn to imitate verbose but uninformative patterns, which can degrade performance. To this end, we incorporate integrated gradient attribution to quantify each token's influence on final answers and aggregate them into two segment-level metrics: (1) attribution strength measures the overall attribution magnitude; and (2) direction consistency captures whether tokens'attributions within a segment are uniformly positive or negative (high consistency), or a mixture of both (moderate consistency). Based on these two metrics, we propose a segment-level selective learning framework to identify important segments with high attribution strength but moderate consistency that indicate reflective rather than shallow reasoning. The framework then applies selective SFT on these important segments while masking loss for unimportant ones. Experiments across multiple models and datasets show that our approach improves accuracy and output efficiency, enabling more effective learning from long reasoning traces .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b13a256c-4e2b-4dce-a8d1-912bcc528c26Cited by top-tier papers1
Ask how each one uses itBuilds on13
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- Solving Quantitative Reasoning Problems with Language ModelsAitor Lewkowycz, Anders Andreassen, David Dohan, Ethan Dyer et al.NeurIPS 2022 · 2,039 citations
- Beyond the 80/20 Rule: High-Entropy Minority Tokens Drive Effective Reinforcement Learning for LLM ReasoningShenzhi Wang, Le Yu, Chang Gao, Chujie Zheng et al.NeurIPS 2025 · 592 citations
- When More is Less: Understanding Chain-of-Thought Length in LLMsYuyang Wu, Yifei Wang, Ziyu Ye, Tianqi Du et al.ICLR 2026 · 225 citations
- Reasoning with Exploration: An Entropy PerspectiveDaixuan Cheng, Shaohan Huang, Xuekai Zhu, Bo Dai et al.AAAI 2026 · 216 citations
Related papers
- Uncovering the Gradient Geometry of Long CoT: A Spectral-guided Approach to Reasoning DistillationSinan Fan, Xiaofeng Sun, Chen Shen, Chenxi Huang et al.ICML 2026
- Your Reasoning Model Knows What Counts: Self-Guided Chain-of-Thought Pruning for Efficient ReasoningZi-Ao Ma, Xian-Ling Mao, Tian Lan, Chen Xu et al.ACL 2026
- When Thinking Fails: The Pitfalls of Reasoning for Instruction-Following in LLMsXiaomin Li, Zhou Yu, Zhiwei Zhang, Xupeng Chen et al.NeurIPS 2025 · 63 citations
- DLoFT: Gradient-Decoupled Fine-Tuning for Generalizable Long Chain-of-Thought ReasoningSitong Wu, Haoru Tan, Jingyao Li, Shaofeng Zhang et al.NeurIPS 2025
- SLAT: Segment-Level Adaptive Trimming for Efficient CoT ReasoningJian Yao, Xiongcai Luo, Ran Cheng, KC TanICML 2026
