Semantics-Adaptive Activation Intervention for LLMs via Dynamic Steering Vectors
Weixuan Wang, Jingyuan Yang, Wei Peng
摘要
Large language models (LLMs) have achieved remarkable performance across many tasks, yet aligning them with desired behaviors remains challenging. Activation intervention has emerged as an effective and economical method to modify the behavior of LLMs. Despite considerable interest in this area, current intervention methods exclusively employ a fixed steering vector to modify model activations, lacking adaptability to diverse input semantics. To address this limitation, we propose Semantics-Adaptive Dynamic Intervention (SADI), a novel method that constructs a dynamic steering vector to intervene model activations at inference time. More specifically, SADI utilizes activation differences in contrastive pairs to precisely identify critical elements of an LLM (i.e., attention heads, hidden states, and neurons) for targeted intervention. During inference, SADI dynamically steers model behavior by scaling element-wise activations based on the directions of input semantics. Experimental results show that SADI outperforms established baselines by substantial margins, improving task performance without training. SADI's cost-effectiveness and generalizability across various LLM backbones and tasks highlight its potential as a versatile alignment technique. 1
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper18
- Angular Steering: Behavior Control via Rotation in Activation SpaceMinh Hieu Vu, Tan M. NguyenNeurIPS 2025 · 被引用 53 次
- ODESteer: A Unified ODE-Based Steering Framework for LLM AlignmentHongjue Zhao, Haosen Sun, Jiangtao Kong, Xiaochang Li 等ICLR 2026 · 被引用 16 次
- Spherical Steering: Geometry-Aware Activation Rotation for Language ModelsZejia You, Chunyuan Deng, Hanjie ChenICML 2026 · 被引用 12 次
- Steering When Necessary: Flexible Steering Large Language Models with BacktrackingZifeng Cheng, Jinwei Gan, Zhiwei Jiang, Cong Wang 等NeurIPS 2025 · 被引用 9 次
- Dynamic Multimodal Activation Steering for Hallucination Mitigation in Large Vision-Language ModelsJianghao Yin, Qin Chen, Kedi Chen, Jie Zhou 等ICLR 2026 · 被引用 7 次
它引用的顶会 Paper18
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou 等ICLR 2021 · 被引用 7,905 次
- Finetuned Language Models are Zero-Shot LearnersJason Wei, Maarten Bosma, Vincent Y. Zhao, Kelvin Guu 等ICLR 2022 · 被引用 4,966 次
- TruthfulQA: Measuring How Models Mimic Human FalsehoodsStephanie Lin, Jacob Hilton, Owain EvansACL 2022 · 被引用 3,228 次
- WinoGrande: An Adversarial Winograd Schema Challenge at ScaleKeisuke Sakaguchi, Ronan Le Bras, Chandra Bhagavatula, Yejin ChoiAAAI 2020 · 被引用 3,037 次
- Inference-Time Intervention: Eliciting Truthful Answers from a Language ModelKenneth Li, Oam Patel, Fernanda B. Viégas, Hanspeter Pfister 等NeurIPS 2023 · 被引用 1,549 次
相关 Paper
- Local Linearity of LLMs Enables Activation Steering via Model-Based Linear Optimal ControlJulian Skifstad, Xinyue Annie Yang, Glen ChouICML 2026 · 被引用 2 次
- Controllable and Explainable Personality Sliders for LLMs at Inference TimeFlorian Hoppe, David Khachaturov, Robert Mullins, Mark Huasong MengICML 2026 · 被引用 1 次
- From Weights to Activations: Is Steering the Next Frontier of Adaptation?Simon Ostermann, Daniil Gurgurov, Tanja Baeumel, Michael A. Hedderich 等ACL 2026 · 被引用 3 次
- Steer Like the LLM: Activation Steering that Mimics PromptingGeert Heyman, Frederik VandeputteICML 2026
- LayerNavigator: Finding Promising Intervention Layers for Efficient Activation Steering in Large Language ModelsHao Sun, Huailiang Peng, Qiong Dai, Xu Bai 等NeurIPS 2025 · 被引用 4 次
