Improving Instruction-Following in Language Models through Activation Steering
Alessandro Stolfo, Vidhisha Balachandran, Safoora Yousefi, Eric Horvitz, Besmira Nushi
Abstract
The ability to follow instructions is crucial for numerous real-world applications of language models. In pursuit of deeper insights and more powerful capabilities, we derive instruction-specific vector representations from language models and use them to steer models accordingly. These vectors are computed as the difference in activations between inputs with and without instructions, enabling a modular approach to activation steering. We demonstrate how this method can enhance model adherence to constraints such as output format, length, and word inclusion, providing inference-time control over instruction following. Our experiments across four models demonstrate how we can use the activation vectors to guide models to follow constraints even without explicit instructions and to enhance performance when instructions are present. Additionally, we explore the compositionality of activation steering, successfully applying multiple instructions simultaneously. Finally, we demonstrate that steering vectors computed on instruction-tuned models can transfer to improve base models. Our findings demonstrate that activation steering offers a practical and scalable approach for fine-grained control in language generation. Our code and data are available at https://github.com/microsoft/llm-steer-instruct .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext dfbbabef-6eba-405c-bc82-afbf1a674eb3Cited by top-tier papers42
- Angular Steering: Behavior Control via Rotation in Activation SpaceMinh Hieu Vu, Tan M. NguyenNeurIPS 2025 · 53 citations
- MCP Security Bench (MSB): Benchmarking Attacks Against Model Context Protocol in LLM AgentsDongsen Zhang, Zekun Li, Xu Luo, Xuannan Liu et al.ICLR 2026 · 47 citations
- Towards High Data Efficiency in Reinforcement Learning with Verifiable RewardXinyu Tang, Zhenduo Zhang, Yurou Liu, Xin Zhao et al.ICLR 2026 · 18 citations
- Transferring Linear Features Across Language Models With Model StitchingAlan Chen, Jack Merullo, Alessandro Stolfo, Ellie PavlickNeurIPS 2025 · 17 citations
- How do LLMs Compute Verbal Confidence?Dharshan Kumaran, Arthur Conmy, Federico Barbero, Simon Osindero et al.ICML 2026 · 16 citations
Builds on36
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- Large Language Models are Zero-Shot ReasonersTakeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo et al.NeurIPS 2022 · 8,168 citations
- Finetuned Language Models are Zero-Shot LearnersJason Wei, Maarten Bosma, Vincent Y. Zhao, Kelvin Guu et al.ICLR 2022 · 4,966 citations
- Multitask Prompted Training Enables Zero-Shot Task GeneralizationVictor Sanh, Albert Webson, Colin Raffel, Stephen H. Bach et al.ICLR 2022 · 1,976 citations
- Inference-Time Intervention: Eliciting Truthful Answers from a Language ModelKenneth Li, Oam Patel, Fernanda B. Viégas, Hanspeter Pfister et al.NeurIPS 2023 · 1,549 citations
Related papers
- Analysing the Generalisation and Reliability of Steering VectorsDaniel Tan, David Chanin, Aengus Lynch, Brooks Paige et al.NeurIPS 2024
- A Simple Yet Effective Method for Non-Refusing Context Relevant Fine-grained Safety Steering in LLMsShaona Ghosh, Amrita Bhattacharjee, Yftah Ziser, Christopher ParisienEMNLP 2025
- Enhancing Instruction Following of LLMs via Activation Steering with Dynamic RejectionMinjae Kang, Jaehyung KimICLR 2026 · 6 citations
- Steer Like the LLM: Activation Steering that Mimics PromptingGeert Heyman, Frederik VandeputteICML 2026
- Steering off Course: Reliability Challenges in Steering Language ModelsPatrick Queiroz Da Silva, Hari Sethuraman, Dheeraj Rajagopal, Hannaneh Hajishirzi et al.ACL 2025 · 17 citations
