Time-Frequency Token Advantage Clipping for Training Efficient Large Reasoning Model
Rong Bao, Bo Wang, Xiao Wang, Hongyu Li, Rui Zheng, Leszek Rutkowski, Qi Zhang, Liang Ding, Dacheng Tao
Abstract
Long Chain-of-Thought (CoT) reasoning enhances large reasoning models' performance but suffers from severe inefficiencies, as models often overthink simple problems or underthink complex ones. Current sequence-level optimizations, like length penalties, are too coarse-grained to distinguish core logic from verbose language, precluding the necessary token-level control for efficient reasoning CoT. To overcome these limitations, we introduce Time-Frequency token Advantage Clipping (TFAC), a novel training framework designed to build efficient large reasoning models via token-level interventions. Specifically, TFAC functions along two dimensions:
- The Frequency Dimension: It discourages inefficient loops and encourages deeper exploration by dynamically reducing the advantage scores of high-entropy tokens that are repeatedly generated within a single reasoning path. 2) The Time Dimension: It reduces excessive overthinking of the system by establishing a historical baseline for the occurrence count of each critical token in previously successful trajectories, and clipping the advantages of tokens that exceed this baseline during training. Crucially, to preserve the model's exploratory capabilities on novel problems, this suppression mechanism is automatically disabled when no historical record of success is available. Experiments conducted on the Deepseek-Distill-32B and Qwen3-8B models show that TFAC outperforms leading baseline methods, improving performance by 2.3 and 3.1 percentage points, respectively, while simultaneously reducing inference costs by 35% and 28% in scenarios where correct answers are generated. These results validate the significant efficacy of TFAC in training large reasoning models that are both powerful and highly efficient.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on14
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- Let's Verify Step by StepHunter Lightman, Vineet Kosaraju, Yuri Burda, Harrison Edwards et al.ICLR 2024 · 3,045 citations
- On the Planning Abilities of Large Language Models - A Critical InvestigationKarthik Valmeekam, Matthew Marquez, Sarath Sreedharan, Subbarao KambhampatiNeurIPS 2023 · 509 citations
- CoT-Valve: Length-Compressible Chain-of-Thought TuningXinyin Ma, Guangnian Wan, Runpeng Yu, Gongfan Fang et al.ACL 2025 · 162 citations
- Thinkless: LLM Learns When to ThinkGongfan Fang, Xinyin Ma, Xinchao WangNeurIPS 2025 · 128 citations
Related papers
- Think Better, Not Longer: Token-Level Marginal Utility for Efficient Reasoning in Large Reasoning ModelsJiawei Li, Yang Gao, Huashan Sun, Chong FengACL 2026
- Anti-Length Shift: Dynamic Outlier Truncation for Training Efficient Reasoning ModelsWei Wu, Liyi Chen, Congxi Xiao, Tianfu Wang et al.ACL 2026 · 1 citation
- Aware First, Think Less: Dynamic Boundary Self-Awareness Drives Significant Gains in Reasoning Efficiency in Large Language ModelsQiguang Chen, Dengyun Peng, Jinhao Liu, Huikang Su et al.AAAI 2026
- Ada-R1: Hybrid-CoT via Bi-Level Adaptive Reasoning OptimizationHaotian Luo, Haiying He, Yibo Wang, Jinluan Yang et al.NeurIPS 2025 · 29 citations
- Overthinking Reduction with Decoupled Rewards and Curriculum Data SchedulingShuyang Jiang, Yusheng Liao, Ya Zhang, Yanfeng Wang et al.ICLR 2026 · 5 citations
