DISCO: Disentangled Communication Steering for Large Language Models
Max Torop, Aria Masoomi, Masih Eskandar, Jennifer G. Dy
Abstract
A variety of recent methods guide large language model outputs via the inferencetime addition of steering vectors to residual-stream or attention-head representations. In contrast, we propose to inject steering vectors directly into the query and value representation spaces within attention heads. We provide evidence that a greater portion of these spaces exhibit high linear discriminability of concepts -a key property motivating the use of steering vectors-than attention head outputs. We analytically characterize the effect of our method, which we term DISentangled COmmunication (DISCO) Steering, on attention head outputs. Our analysis reveals that DISCO disentangles a strong but underutilized baseline, steering attention head inputs, which implicitly modifies queries and values in a rigid manner. In contrast, DISCO's direct modulation of these components enables more granular control. We find that DISCO achieves superior performance over a number of steering vector baselines across multiple datasets on LLaMA 3.1 8B and Gemma 2 9B, with steering efficacy scoring up to 19.1% higher than the runner-up. Our results support the conclusion that the query and value spaces are powerful building blocks for steering vector methods. Our code is publicly available at https://github.com/MaxTorop/DISCO.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b2a1da5a-b966-417f-995f-8cc6824c12c9Cited by top-tier papers1
Ask how each one uses itBuilds on24
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Locating and Editing Factual Associations in GPTKevin Meng, David Bau, Alex Andonian, Yonatan BelinkovNeurIPS 2022 · 3,415 citations
- TruthfulQA: Measuring How Models Mimic Human FalsehoodsStephanie Lin, Jacob Hilton, Owain EvansACL 2022 · 3,228 citations
- Implicit In-context LearningZhuowei Li, Zihao Xu, Ligong Han, Yunhe Gao et al.ICLR 2025 · 1,989 citations
- Inference-Time Intervention: Eliciting Truthful Answers from a Language ModelKenneth Li, Oam Patel, Fernanda B. Viégas, Hanspeter Pfister et al.NeurIPS 2023 · 1,549 citations
Related papers
- Steering Information Utility in Key-Value Memory for Language Model Post-TrainingChunyuan Deng, Ruidi Chang, Hanjie ChenNeurIPS 2025 · 2 citations
- Endogenous Resistance to Activation Steering in Language ModelsAlex McKenzie, Keenan Pepper, Stijn Servaes, Martin Leitgab et al.ICML 2026 · 3 citations
- Does Higher Interpretability Imply Better Utility? A Pairwise Analysis on Sparse AutoencodersXu Wang, Yan Hu, Benyou Wang, Difan ZouICLR 2026 · 9 citations
- Improved Representation Steering for Language ModelsZhengxuan Wu, Qinan Yu, Aryaman Arora, Christopher D. Manning et al.NeurIPS 2025 · 22 citations
- Compositional Steering of Large Language Models with Steering TokensGorjan Radevski, Kiril Gashteovski, Giwon Hong, Carolin Lawrence et al.ACL 2026 · 4 citations
