Why Steering Works: Toward a Unified View of Language Model Parameter Dynamics
Ziwen Xu, Chenyan Wu, Hengyu Sun, Haiwen Hong, Mengru Wang, Yunzhi Yao, Longtao Huang, Hui Xue, Shumin Deng, Zhixuan Chu, Huajun Chen, Ningyu Zhang
摘要
Methods for controlling large language models (LLMs), including local weight fine-tuning, LoRA-based adaptation, and activation-based interventions, are often studied in isolation, obscuring their connections and making comparison difficult. In this work, we present a unified view that frames these interventions as dynamic weight updates induced by a control signal, placing them within a single conceptual framework. Building on this view, we propose a unified preference-utility analysis that separates control effects into preference, defined as the tendency toward a target concept, and utility, defined as coherent and task-valid generation, and measures both on a shared log-odds scale using polarity-paired contrastive examples. Across methods, we observe a consistent trade-off between preference and utility: stronger control increases preference while predictably reducing utility. We further explain this behavior through an activation manifold perspective, in which control shifts representations along target-concept directions to enhance preference, while utility declines primarily when interventions push representations off the model's valid-generation manifold. Finally, we introduce a new steering approach SPLIT guided by this analysis that improves preference while better preserving utility. Code is available at https://github.com/zjunlp/EasyEdit/blob/main/examples/SPLIT.md.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Stop When Further Reasoning Won’t Help: Attention-State Adaptive Generation in Reasoning ModelsJiakai Li, KE QIN, Rongzheng Wang, Yizhuo Ma 等ICML 2026
- How Controllable Are Large Language Models? A Unified Evaluation across Behavioral GranularitiesZiwen Xu, Kewei Xu, Haoming Xu, Haiwen Hong 等ACL 2026
它引用的顶会 Paper18
- The Linear Representation Hypothesis and the Geometry of Large Language ModelsKiho Park, Yo Joong Choe, Victor VeitchICML 2024 · 被引用 461 次
- LoRA+: Efficient Low Rank Adaptation of Large ModelsSoufiane Hayou, Nikhil Ghosh, Bin YuICML 2024 · 被引用 388 次
- VeRA: Vector-based Random Matrix AdaptationDawid Jan Kopiczko, Tijmen Blankevoort, Yuki M. AsanoICLR 2024 · 被引用 308 次
- Personalized Steering of Large Language Models: Versatile Steering Vectors Through Bi-directional Preference OptimizationYuanpu Cao, Tianrong Zhang, Bochuan Cao, Ziyi Yin 等NeurIPS 2024 · 被引用 135 次
- mHC: Manifold-Constrained Hyper-ConnectionsZhenda Xie, Yixuan Wei, Huanqi Cao, Chenggang Zhao 等ICML 2026 · 被引用 72 次
相关 Paper
- From Weights to Activations: Is Steering the Next Frontier of Adaptation?Simon Ostermann, Daniil Gurgurov, Tanja Baeumel, Michael A. Hedderich 等ACL 2026 · 被引用 3 次
- Causal-Steer: Disentangled Continuous Style Control without Parallel CorporaQingsong Wang, Chang Yao, Jingyuan ChenICLR 2026
- Steering Language Models with Weight ArithmeticConstanza Fierro, Fabien RogerICLR 2026 · 被引用 19 次
- Belief Dynamics Reveal the Dual Nature of In-Context Learning and Activation SteeringEric Bigelow, Daniel Wurgaft, YingQiao Wang, Noah Goodman 等ICML 2026
- Towards a Unified Paradigm of Concept Editing in Large Language ModelsZhuowen Han, Xinwei Wu, Dan Shi, Renren Jin 等EMNLP 2025
