Steering off Course: Reliability Challenges in Steering Language Models
Patrick Queiroz Da Silva, Hari Sethuraman, Dheeraj Rajagopal, Hannaneh Hajishirzi, Sachin Kumar
摘要
Steering methods for language models (LMs) have gained traction as lightweight alternatives to fine-tuning, enabling targeted modifications to model activations. However, prior studies primarily report results on a few models, leaving critical gaps in understanding the robustness of these methods. In this work, we systematically examine three prominent steering methods -- DoLa, function vectors, and task vectors. In contrast to the original studies, which evaluated a handful of models, we test up to 36 models belonging to 14 families with sizes ranging from 1.5B to 70B parameters. Our experiments reveal substantial variability in the effectiveness of the steering approaches, with a large number of models showing no improvement and at times degradation in steering performance. Our analysis demonstrate fundamental flaws in the assumptions underlying these methods, challenging their reliability as scalable steering solutions.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- IF-Guide: Influence Function-Guided Detoxification of LLMsZachary Coalson, Juhan Bae, Nicholas Carlini, Sanghyun HongNeurIPS 2025 · 被引用 9 次
- Compositional Steering of Large Language Models with Steering TokensGorjan Radevski, Kiril Gashteovski, Giwon Hong, Carolin Lawrence 等ACL 2026 · 被引用 4 次
- IA2: Alignment with ICL Activations improves Supervised Fine-TuningAayush Mishra, Daniel Khashabi, Anqi LiuICLR 2026 · 被引用 1 次
- Eliciting Chain-of-Thought in Base LLMs via Gradient-Based Representation OptimizationZijian Wang, Yanxiang Ma, Chang XuAAAI 2026
- ThoughtProbe: Classifier-Guided LLM Thought Space Exploration via Probing RepresentationsZijian Wang, Chang XuEMNLP 2025
它引用的顶会 Paper18
- Locating and Editing Factual Associations in GPTKevin Meng, David Bau, Alex Andonian, Yonatan BelinkovNeurIPS 2022 · 被引用 3,415 次
- TruthfulQA: Measuring How Models Mimic Human FalsehoodsStephanie Lin, Jacob Hilton, Owain EvansACL 2022 · 被引用 3,228 次
- Inference-Time Intervention: Eliciting Truthful Answers from a Language ModelKenneth Li, Oam Patel, Fernanda B. Viégas, Hanspeter Pfister 等NeurIPS 2023 · 被引用 1,549 次
- Sparse Autoencoders Find Highly Interpretable Features in Language ModelsRobert Huben, Hoagy Cunningham, Logan Riggs Smith, Aidan Ewart 等ICLR 2024 · 被引用 1,072 次
- Towards Automated Circuit Discovery for Mechanistic InterpretabilityArthur Conmy, Augustine N. Mavor-Parker, Aengus Lynch, Stefan Heimersheim 等NeurIPS 2023 · 被引用 861 次
相关 Paper
- Analysing the Generalisation and Reliability of Steering VectorsDaniel Tan, David Chanin, Aengus Lynch, Brooks Paige 等NeurIPS 2024
- From Weights to Activations: Is Steering the Next Frontier of Adaptation?Simon Ostermann, Daniil Gurgurov, Tanja Baeumel, Michael A. Hedderich 等ACL 2026 · 被引用 3 次
- Steer Like the LLM: Activation Steering that Mimics PromptingGeert Heyman, Frederik VandeputteICML 2026
- GrAInS: Gradient-based Attribution for Inference-Time Steering of LLMs and VLMsDuy Nguyen, Archiki Prasad, Elias Stengel-Eskin, Mohit BansalACL 2026 · 被引用 4 次
- Improving Instruction-Following in Language Models through Activation SteeringAlessandro Stolfo, Vidhisha Balachandran, Safoora Yousefi, Eric Horvitz 等ICLR 2025
