SAE-SSV: Supervised Steering in Sparse Representation Spaces for Reliable Control of Language Models
Zirui He, Mingyu Jin, Bo Shen, Ali Payani, Yongfeng Zhang, Mengnan Du
Abstract
Large language models (LLMs) have demonstrated impressive capabilities in natural language understanding and generation, but controlling their behavior reliably remains challenging, especially in open-ended generation settings. This paper introduces a novel supervised steering approach that operates in sparse, interpretable representation spaces. We employ sparse autoencoders (SAEs) to obtain sparse latent representations that aim to disentangle semantic attributes from model activations. Then we train linear classifiers to identify a small subspace of task-relevant dimensions in latent representations. Finally, we learn supervised steering vectors constrained to this subspace, optimized to align with target behaviors. Experiments across sentiment, truthfulness, and politics polarity steering tasks with multiple LLMs demonstrate that our supervised steering vectors achieve higher success rates with minimal degradation in generation quality compared to existing methods. Further analysis reveals that a notably small subspace is sufficient for effective steering, enabling more targeted and interpretable interventions. Our implementation is publicly available at https: //github.com/Ineedanamehere/SAE-SSV .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext ee7b612b-a9fa-4ce9-85f3-52e23273f261Cited by top-tier papers5
- All Circuits Lead to Rome: Rethinking Functional Anisotropy in Circuit and Sheaf Discovery for LLMsXi Chen, Mingyu Jin, Jingcheng (Frank) Niu, Yutong Yin et al.ICML 2026 · 5 citations
- Exploring Diverse Generation Paths via Inference-time Stiefel Activation SteeringDongxuan Zhu, Ly Tran Ho Khanh, Andy Yat-Ming Cheung, Man-Chung Yue et al.ICLR 2026 · 4 citations
- Where Concept Erasure Should Occur: Concept–Layer Alignment in Text-to-Video Diffusion ModelsYiwei Xie, Ping Liu, Zheng ZhangICML 2026
- PrivSV: Differentially Private Steering Vector for Large Language ModelsHaocheng Yang, Xiang Cheng, Chenhao Sun, Pengfei Zhang et al.AAAI 2026
- SDA: Steering-Driven Distribution Alignment for Open LLMs Without Fine-TuningWei Xia, Zhi-Hong DengAAAI 2026
Builds on14
- Finetuned Language Models are Zero-Shot LearnersJason Wei, Maarten Bosma, Vincent Y. Zhao, Kelvin Guu et al.ICLR 2022 · 4,966 citations
- Inference-Time Intervention: Eliciting Truthful Answers from a Language ModelKenneth Li, Oam Patel, Fernanda B. Viégas, Hanspeter Pfister et al.NeurIPS 2023 · 1,549 citations
- Plug and Play Language Models: A Simple Approach to Controlled Text GenerationSumanth Dathathri, Andrea Madotto, Janice Lan, Jane Hung et al.ICLR 2020 · 1,166 citations
- Refusal in Language Models Is Mediated by a Single DirectionAndy Arditi, Oscar Obeso, Aaquib Syed, Daniel Paleka et al.NeurIPS 2024 · 1,166 citations
- Sparse Autoencoders Find Highly Interpretable Features in Language ModelsRobert Huben, Hoagy Cunningham, Logan Riggs Smith, Aidan Ewart et al.ICLR 2024 · 1,072 citations
Related papers
- Unveiling Language-Specific Features in Large Language Models via Sparse AutoencodersBoyi Deng, Yu Wan, Baosong Yang, Yidan Zhang et al.ACL 2025
- Measuring and Guiding MonosemanticityRuben Härle, Felix Friedrich, Manuel Brack, Björn Deiseroth et al.NeurIPS 2025 · 12 citations
- Sparse Autoencoders for Interpretable Emotion Control in Text-to-SpeechHongfei Du, Jiacheng Shi, Sidi Lu, Gang Zhou et al.ICML 2026
- Does Higher Interpretability Imply Better Utility? A Pairwise Analysis on Sparse AutoencodersXu Wang, Yan Hu, Benyou Wang, Difan ZouICLR 2026 · 9 citations
- Uncovering Sentiment Analysis Circuit in Large Language ModelShichen Li, Zhouyang Wang, Zhongqing Wang, Peifeng LiACL 2026
