Controllable LLM Reasoning via Sparse Autoencoder-Based Steering
Yi Fang, Wenjie Wang, Mingfeng Xue, Boyi Deng, Fengli Xu, Dayiheng Liu, Fuli Feng
摘要
Large Reasoning Models (LRMs) exhibit human-like cognitive reasoning strategies (e.g. backtracking, cross-verification) during reasoning process, which improves their performance on complex tasks. Currently, reasoning strategies are autonomously selected by LRMs themselves. However, such autonomous selection often produces inefficient or even erroneous reasoning paths. To make reasoning more reliable and flexible, it is important to develop methods for controlling reasoning strategies. Existing methods struggle to control fine-grained reasoning strategies due to conceptual entanglement in LRMs'hidden states. To address this, we leverage Sparse Autoencoders (SAEs) to decompose strategy-entangled hidden states into a disentangled feature space. To identify the few strategy-specific features from the vast pool of SAE features, we propose SAE-Steering, an efficient two-stage feature identification pipeline. SAE-Steering first recalls features that amplify the logits of strategy-specific keywords, filtering out over 99% of features, and then ranks the remaining features by their control effectiveness. Using the identified strategy-specific features as control vectors, SAE-Steering outperforms existing methods by over 15% in control effectiveness. Furthermore, controlling reasoning strategies can redirect LRMs from erroneous paths to correct ones, achieving a 7% absolute accuracy improvement.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Do Sparse Autoencoders Identify Reasoning Features in Language Models?George Ma, Zhongyuan Liang, Irene Y. Chen, Somayeh SojoudiICML 2026 · 被引用 9 次
- SASFT: Sparse Autoencoder-guided Supervised Finetuning to Mitigate Unexpected Code-Switching in LLMsBoyi Deng, Yu Wan, Baosong Yang, Fei Huang 等ICLR 2026 · 被引用 2 次
- The Tell-Tale Norm: Magnitude as a Signal for Reasoning Dynamics in Large Language ModelsJinyang Zhang, Hongxin Ding, Yue Fang, Weibin Liao 等ICML 2026
它引用的顶会 Paper15
- Reflexion: language agents with verbal reinforcement learningNoah Shinn, Federico Cassano, Ashwin Gopinath, Karthik Narasimhan 等NeurIPS 2023 · 被引用 5,828 次
- Sparse Autoencoders Find Highly Interpretable Features in Language ModelsRobert Huben, Hoagy Cunningham, Logan Riggs Smith, Aidan Ewart 等ICLR 2024 · 被引用 1,072 次
- LMSYS-Chat-1M: A Large-Scale Real-World LLM Conversation DatasetLianmin Zheng, Wei-Lin Chiang, Ying Sheng, Tianle Li 等ICLR 2024 · 被引用 419 次
- SELF-DISCOVER: Large Language Models Self-Compose Reasoning StructuresPei Zhou, Jay Pujara, Xiang Ren, Xinyun Chen 等NeurIPS 2024 · 被引用 151 次
- Dense Passage Retrieval for Open-Domain Question AnsweringVladimir Karpukhin, Barlas Oguz, Sewon Min, Patrick Lewis 等EMNLP 2020 · 被引用 142 次
相关 Paper
- Step-Level Sparse Autoencoder for Reasoning Process InterpretationXuan Yang, Jiayu Liu, Yuhang Lai, Hao Xu 等ICML 2026 · 被引用 2 次
- Feature Extraction and Steering for Enhanced Chain-of-Thought Reasoning in Language ModelsZihao Li, Xu Wang, Yuzhe Yang, Ziyu Yao 等EMNLP 2025 · 被引用 14 次
- Fantastic Reasoning Behaviors and Where to Find Them: Unsupervised Discovery of the Reasoning ProcessZhenyu Zhang, Shujian Zhang, John Lambert, Wenxuan Zhou 等ICML 2026 · 被引用 3 次
- ActivationReasoning: Logical Reasoning in Latent Activation SpacesLukas Helff, Ruben Härle, Wolfgang Stammer, Felix Friedrich 等ICLR 2026 · 被引用 6 次
- Beyond Prompt Engineering: Robust Behavior Control in LLMs via Steering Target AtomsMengru Wang, Ziwen Xu, Shengyu Mao, Shumin Deng 等ACL 2025 · 被引用 19 次
