Multi-Attribute Steering of Language Models via Targeted Intervention
Duy Nguyen, Archiki Prasad, Elias Stengel-Eskin, Mohit Bansal
Abstract
Inference-time intervention (ITI) has emerged as a promising method for steering large language model (LLM) behavior in a particular direction (e.g., improving helpfulness) by intervening on token representations without costly updates to the LLM's parameters. However, existing ITI approaches fail to scale to multi-attribute settings with conflicts, such as enhancing helpfulness while also reducing toxicity. To address this, we introduce Multi-Attribute Targeted Steering (MAT-Steer), a novel steering framework designed for selective token-level intervention across multiple attributes. MAT-Steer learns steering vectors using an alignment objective that shifts the model's internal representations of undesirable outputs closer to those of desirable ones while enforcing sparsity and orthogonality among vectors for different attributes, thereby reducing inter-attribute conflicts. We evaluate MAT-Steer in two distinct settings: (i) on question answering (QA) tasks where we balance attributes like truthfulness, bias, and toxicity; (ii) on generative tasks where we simultaneously improve attributes like helpfulness, correctness, and coherence. MAT-Steer outperforms existing ITI and parameter-efficient fine-tuning approaches across both task types (e.g., 3% average accuracy gain across QA tasks and 55.82% win rate against the best ITI baseline).
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers7
- ODESteer: A Unified ODE-Based Steering Framework for LLM AlignmentHongjue Zhao, Haosen Sun, Jiangtao Kong, Xiaochang Li et al.ICLR 2026 · 16 citations
- Compositional Steering of Large Language Models with Steering TokensGorjan Radevski, Kiril Gashteovski, Giwon Hong, Carolin Lawrence et al.ACL 2026 · 4 citations
- From Weights to Activations: Is Steering the Next Frontier of Adaptation?Simon Ostermann, Daniil Gurgurov, Tanja Baeumel, Michael A. Hedderich et al.ACL 2026 · 3 citations
- StyliTruth : Unlocking Stylized yet Truthful LLM Generation via Disentangled SteeringChenglei Shen, Zhongxiang Sun, Teng Shi, Xiao Zhang et al.ICLR 2026
- Knowledge Vector of Logical Reasoning in Large Language ModelsZixuan Wang, Yuanyuan LeiACL 2026
Builds on23
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning et al.NeurIPS 2023 · 10,924 citations
- TruthfulQA: Measuring How Models Mimic Human FalsehoodsStephanie Lin, Jacob Hilton, Owain EvansACL 2022 · 3,228 citations
- Inference-Time Intervention: Eliciting Truthful Answers from a Language ModelKenneth Li, Oam Patel, Fernanda B. Viégas, Hanspeter Pfister et al.NeurIPS 2023 · 1,549 citations
Related papers
- FineSteer: A Unified Framework for Fine-Grained Inference-Time Steering in Large Language ModelsZixuan Weng, Jinghuai Zhang, Kunlin Cai, Ying Li et al.ACL 2026
- REAL: Reading Out Transformer Activations for Precise Localization in Language Model SteeringLi-Ming Zhan, Bo LIU, Yujie Feng, Chengqiang Xie et al.ICLR 2026 · 4 citations
- GrAInS: Gradient-based Attribution for Inference-Time Steering of LLMs and VLMsDuy Nguyen, Archiki Prasad, Elias Stengel-Eskin, Mohit BansalACL 2026 · 4 citations
- SAE-SSV: Supervised Steering in Sparse Representation Spaces for Reliable Control of Language ModelsZirui He, Mingyu Jin, Bo Shen, Ali Payani et al.EMNLP 2025 · 1 citation
- Steering When Necessary: Flexible Steering Large Language Models with BacktrackingZifeng Cheng, Jinwei Gan, Zhiwei Jiang, Cong Wang et al.NeurIPS 2025 · 9 citations
