CATS: Category-Aware Token-level Steering for Training-Free Redundancy Reduction in Large Reasoning Models
Mengfei Zhang, Zhenglin Wang
摘要
While Large Reasoning Models (LRMs) exhibit remarkable capabilities in complex tasks, they often suffer from excessive redundancy in their chain-of-thought reasoning. This significantly reduces inference efficiency and increases computational costs. We identify that LRM redundancy is not uniformly homogeneous but can be taxonomized according to whether it is destructive to the final answer: destructive redundancy (e.g., logical drift, hallucination amplification) versus non-destructive redundancy (e.g., repetition, over-elaboration). Moreover, LRM's redundant and concise responses exhibit a significant distinction in their hidden layer representation spaces. Based on these insights, we propose CATS (Category-Aware Token-level Steering), a training-free and lightweight method to reduce the redundancy phenomenon. CATS decomposes redundancy into six semantically interpretable characteristic dimensions. By flexibly weighting and combining the differential vectors corresponding to these dimensions, CATS synthesizes a composite intervention vector, enabling zero-parameter intervention in the hidden layers. Experiments across three LRM models and five mathematical reasoning datasets demonstrate that CATS reduces reasoning length by an average of 25% while maintaining or even slightly improving task accuracy. CATS offers a pluggable, training-free, and lightweight solution, making it particularly beneficial for users in low-resource environments.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper3
- Let's Verify Step by StepHunter Lightman, Vineet Kosaraju, Yuri Burda, Harrison Edwards 等ICLR 2024 · 被引用 3,045 次
- C3oT: Generating Shorter Chain-of-Thought Without Compromising EffectivenessYu Kang, Xianghui Sun, Liangyu Chen, Wei ZouAAAI 2025 · 被引用 162 次
- TokenSkip: Controllable Chain-of-Thought Compression in LLMsHeming Xia, Chak Tou Leong, Wenjie Wang, Yongqi Li 等EMNLP 2025 · 被引用 6 次
相关 Paper
- SLAT: Segment-Level Adaptive Trimming for Efficient CoT ReasoningJian Yao, Xiongcai Luo, Ran Cheng, KC TanICML 2026
- Efficient Reasoning with Balanced ThinkingYulin Li, Tengyao Tu, Li Ding, Junjie Wang 等ICLR 2026 · 被引用 7 次
- Adaptive Spatial and Temporal Redundancy Optimization for Efficient Reasoning in Large Language ModelsTianle Chen, Pengyu Cheng, Qiyuan Zhu, Jiacheng Wang 等ACL 2026 · 被引用 1 次
- SAT: Balancing Reasoning Accuracy and Efficiency with Stepwise Adaptive ThinkingWeiyang Huang, Xuefeng Bai, Kehai Chen, Xinyang Chen 等ACL 2026 · 被引用 3 次
- Unlocking the Black Box of Latent Reasoning: An Interpretability-Guided Approach to InterventionShuochen Chang, Tong Bai, Xiaofeng Zhang, Qianli Ma 等ACL 2026 · 被引用 1 次
