LayerNavigator: Finding Promising Intervention Layers for Efficient Activation Steering in Large Language Models
Hao Sun, Huailiang Peng, Qiong Dai, Xu Bai, Yanan Cao
摘要
Activation steering is an efficient technique for aligning the behavior of large language models (LLMs) by injecting steering vectors directly into a model's residual stream during inference. A pivotal challenge in this approach lies in choosing the right layers to intervene, as inappropriate selection can undermine behavioral alignment and even impair the model's language fluency and other core capabilities. While single-layer steering allows straightforward evaluation on held-out data to identify the "best" layer, it offers only limited alignment improvements. Multi-layer steering promises stronger control but faces a combinatorial explosion of possible layer subsets, making exhaustive search impractical. To address these challenges, we propose LayerNavigator, which provides a principled and promising layer selection strategy. The core innovation of LayerNavigator lies in its novel, quantifiable criterion that evaluates each layer's steerability by jointly considering two key aspects: discriminability and consistency. By reusing the activations computed during steering vector generation, LayerNavigator requires no extra data and adds negligible overhead. Comprehensive experiments show that LayerNavigator achieves not only superior alignment but also greater scalability and interpretability compared to existing strategies. Our code is available at https://github.com/Bryson-Arrot/LayerNavigator
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- DLM-Scope: Mechanistic Interpretability of Diffusion Language Models via Sparse AutoencodersXu Wang, Bingqing Jiang, Yu Wan, Baosong Yang 等ICML 2026
- Towards Steering without Sacrifice: Principled Training of Steering Vectors for Prompt-only InterventionsYuntai Bao, Qinfeng Li, Xinyan Yu, Ge Su 等ICML 2026
它引用的顶会 Paper14
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning 等NeurIPS 2023 · 被引用 10,924 次
- Inference-Time Intervention: Eliciting Truthful Answers from a Language ModelKenneth Li, Oam Patel, Fernanda B. Viégas, Hanspeter Pfister 等NeurIPS 2023 · 被引用 1,549 次
- Refusal in Language Models Is Mediated by a Single DirectionAndy Arditi, Oscar Obeso, Aaquib Syed, Daniel Paleka 等NeurIPS 2024 · 被引用 1,166 次
- Investigating Gender Bias in Language Models Using Causal Mediation AnalysisJesse Vig, Sebastian Gehrmann, Yonatan Belinkov, Sharon Qian 等NeurIPS 2020 · 被引用 851 次
相关 Paper
- Analysing the Generalisation and Reliability of Steering VectorsDaniel Tan, David Chanin, Aengus Lynch, Brooks Paige 等NeurIPS 2024
- SCANS: Mitigating the Exaggerated Safety for LLMs via Safety-Conscious Activation SteeringZouying Cao, Yifei Yang, Hai ZhaoAAAI 2025 · 被引用 35 次
- Controllable and Explainable Personality Sliders for LLMs at Inference TimeFlorian Hoppe, David Khachaturov, Robert Mullins, Mark Huasong MengICML 2026 · 被引用 1 次
- Local Linearity of LLMs Enables Activation Steering via Model-Based Linear Optimal ControlJulian Skifstad, Xinyue Annie Yang, Glen ChouICML 2026 · 被引用 2 次
- Activation Steering with a Feedback ControllerDung Viet Nguyen, Yen Nhi Pham, Hieu M. Vu, Lei Zhang 等ICLR 2026 · 被引用 13 次
