Steering at the Source: Style Modulation Heads for Robust Persona Control
Yoshihiro Izawa, Gouki Minegishi, Koshi Eguchi, Sosuke Hosokawa, Kenjiro Taura
Abstract
Activation steering offers a computationally efficient mechanism for controlling Large Language Models (LLMs) without fine-tuning. While effectively controlling target traits (e.g., persona), coherency degradation remains a major obstacle to safety and practical deployment. We hypothesize that this degradation stems from intervening on the residual stream, which indiscriminately affects aggregated features and inadvertently amplifies off-target noise. In this work, we identify a sparse subset of attention heads (only three heads) that independently govern persona and style formation, which we term Style Modulation Heads . Specifically, these heads can be localized via geometric analysis of internal representations, combining layer-wise cosine similarity and head-wise contribution scores. We demonstrate that intervention targeting only these specific heads achieves robust behavioral control while significantly mitigating the coherency degradation observed in residual stream steering. More broadly, our findings show that precise, component-level localization enables safer and more precise model control.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 4b55dbc4-6fdc-406b-8232-db6c4507217bBuilds on24
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou et al.ICLR 2021 · 7,905 citations
- Locating and Editing Factual Associations in GPTKevin Meng, David Bau, Alex Andonian, Yonatan BelinkovNeurIPS 2022 · 3,415 citations
- Inference-Time Intervention: Eliciting Truthful Answers from a Language ModelKenneth Li, Oam Patel, Fernanda B. Viégas, Hanspeter Pfister et al.NeurIPS 2023 · 1,549 citations
- LLM Evaluators Recognize and Favor Their Own GenerationsArjun Panickssery, Samuel R. Bowman, Shi FengNeurIPS 2024 · 865 citations
- The Linear Representation Hypothesis and the Geometry of Large Language ModelsKiho Park, Yo Joong Choe, Victor VeitchICML 2024 · 461 citations
Related papers
- Controllable and Explainable Personality Sliders for LLMs at Inference TimeFlorian Hoppe, David Khachaturov, Robert Mullins, Mark Huasong MengICML 2026 · 1 citation
- Causal-Steer: Disentangled Continuous Style Control without Parallel CorporaQingsong Wang, Chang Yao, Jingyuan ChenICLR 2026
- Fine-Grained Activation Steering: Steering Less, Achieving MoreZijian Feng, Tianjiao Li, Zixiao Zhu, Hanzhang Zhou et al.ICLR 2026 · 6 citations
- Angular Steering: Behavior Control via Rotation in Activation SpaceMinh Hieu Vu, Tan M. NguyenNeurIPS 2025 · 53 citations
- LayerNavigator: Finding Promising Intervention Layers for Efficient Activation Steering in Large Language ModelsHao Sun, Huailiang Peng, Qiong Dai, Xu Bai et al.NeurIPS 2025 · 4 citations
