Improved Representation Steering for Language Models
Zhengxuan Wu, Qinan Yu, Aryaman Arora, Christopher D. Manning, Christopher Potts
Abstract
Steering methods for language models (LMs) seek to provide fine-grained and interpretable control over model generations by variously changing model inputs, weights, or representations to adjust behavior. Recent work has shown that adjusting weights or representations is often less effective than steering by prompting, for instance when wanting to introduce or suppress a particular concept. We demonstrate how to improve representation steering via our new Reference-free Preference Steering (RePS), a bidirectional preference-optimization objective that jointly does concept steering and suppression. We train three parameterizations of RePS and evaluate them on AXBENCH, a large-scale model steering benchmark. On Gemma models with sizes ranging from 2B to 27B, RePS outperforms all existing steering methods trained with a language modeling objective and substantially narrows the gap with prompting -while promoting interpretability and minimizing parameter count. In suppression, RePS matches the language-modeling objective on Gemma-2 and outperforms it on the larger Gemma-3 variants while remaining resilient to prompt-based jailbreaking attacks that defeat prompting. Overall, our results suggest that RePS provides an interpretable and robust alternative to prompting for both steering and suppression. github.com/stanfordnlp/axbench * Equal contribution. Preprint. Under review.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b49a9b87-4f8a-4718-9167-d01ee25546e4Cited by top-tier papers8
- Spherical Steering: Geometry-Aware Activation Rotation for Language ModelsZejia You, Chunyuan Deng, Hanjie ChenICML 2026 · 12 citations
- Compositional Steering of Large Language Models with Steering TokensGorjan Radevski, Kiril Gashteovski, Giwon Hong, Carolin Lawrence et al.ACL 2026 · 4 citations
- Why Steering Works: Toward a Unified View of Language Model Parameter DynamicsZiwen Xu, Chenyan Wu, Hengyu Sun, Haiwen Hong et al.ACL 2026 · 4 citations
- Faithful Bi-Directional Model Steering via Distribution Matching and Distributed Interchange InterventionsYuntai Bao, Xuhong Zhang, Jintao Chen, Ge Su et al.ICLR 2026 · 2 citations
- Where Paths Split: Localized, Calibrated Control of Moral Reasoning in Large Language ModelsChenchen Yuan, Zheyu Zhang, Gjergji KasneciACL 2026
Builds on23
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning et al.NeurIPS 2023 · 10,924 citations
- Inference-Time Intervention: Eliciting Truthful Answers from a Language ModelKenneth Li, Oam Patel, Fernanda B. Viégas, Hanspeter Pfister et al.NeurIPS 2023 · 1,549 citations
- SimPO: Simple Preference Optimization with a Reference-Free RewardYu Meng, Mengzhou Xia, Danqi ChenNeurIPS 2024 · 1,203 citations
- Towards a Unified View of Parameter-Efficient Transfer LearningJunxian He, Chunting Zhou, Xuezhe Ma, Taylor Berg-Kirkpatrick et al.ICLR 2022 · 1,182 citations
Related papers
- AxBench: Steering LLMs? Even Simple Baselines Outperform Sparse AutoencodersZhengxuan Wu, Aryaman Arora, Atticus Geiger, Zheng Wang et al.ICML 2025
- Steer Like the LLM: Activation Steering that Mimics PromptingGeert Heyman, Frederik VandeputteICML 2026
- Unsupervised Concept Vector Extraction for Bias Control in LLMsHannah Cyberey, Yangfeng Ji, David EvansEMNLP 2025 · 4 citations
- COLD-Steer: Steering Large Language Models via In-Context One-step Learning DynamicsKartik Sharma, Rakshit S. TrivediICLR 2026 · 8 citations
- Concept Heterogeneity-aware Representation SteeringLaziz Abdullaev, Noelle Y. L. Wong, Ryan Lee, Shiqi Jiang et al.ICML 2026
