FineSteer: A Unified Framework for Fine-Grained Inference-Time Steering in Large Language Models
Zixuan Weng, Jinghuai Zhang, Kunlin Cai, Ying Li, Peiran Wang, Yuan Tian
摘要
Large language models (LLMs) often exhibit undesirable behaviors, such as safety violations and hallucinations. Although inference-time steering offers a cost-effective way to adjust model behavior without updating its parameters, existing methods often fail to be simultaneously effective, utilitypreserving, and training-efficient due to their rigid, one-size-fits-all designs and limited adaptability. In this work, we present FineSteer, a novel steering framework that decomposes inference-time steering into two complementary stages-conditional steering and fine-grained vector synthesis-allowing finegrained control over when and how to steer internal representations. In the first stage, we introduce a Subspace-guided Conditional Steering (SCS) mechanism that preserves model utility by avoiding unnecessary steering. In the second stage, we propose a Mixture-of-Steering-Experts (MoSE) mechanism that captures the multimodal nature of desired steering behaviors and generates query-specific steering vectors for improved effectiveness. Through tailored designs in both SCS and MoSE, FineSteer maintains robust performance on general queries while adaptively optimizing steering vectors for targeted inputs in a training-efficient manner. Extensive experiments on safety and truthfulness benchmarks show that FineSteer outperforms the state-of-the-art methods in overall performance (e.g., A 7.6% improvement on TruthfulQA over Llama-3.), achieving stronger steering performance with minimal utility loss. The code is available at https://github.com/YukinoAsuna/FineSteer .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper24
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning 等NeurIPS 2023 · 被引用 10,924 次
- TruthfulQA: Measuring How Models Mimic Human FalsehoodsStephanie Lin, Jacob Hilton, Owain EvansACL 2022 · 被引用 3,228 次
- Jailbroken: How Does LLM Safety Training Fail?Alexander Wei, Nika Haghtalab, Jacob SteinhardtNeurIPS 2023 · 被引用 2,230 次
- Inference-Time Intervention: Eliciting Truthful Answers from a Language ModelKenneth Li, Oam Patel, Fernanda B. Viégas, Hanspeter Pfister 等NeurIPS 2023 · 被引用 1,549 次
- Refusal in Language Models Is Mediated by a Single DirectionAndy Arditi, Oscar Obeso, Aaquib Syed, Daniel Paleka 等NeurIPS 2024 · 被引用 1,166 次
相关 Paper
- A Simple Yet Effective Method for Non-Refusing Context Relevant Fine-grained Safety Steering in LLMsShaona Ghosh, Amrita Bhattacharjee, Yftah Ziser, Christopher ParisienEMNLP 2025
- GrAInS: Gradient-based Attribution for Inference-Time Steering of LLMs and VLMsDuy Nguyen, Archiki Prasad, Elias Stengel-Eskin, Mohit BansalACL 2026 · 被引用 4 次
- Steering When Necessary: Flexible Steering Large Language Models with BacktrackingZifeng Cheng, Jinwei Gan, Zhiwei Jiang, Cong Wang 等NeurIPS 2025 · 被引用 9 次
- Multi-Attribute Steering of Language Models via Targeted InterventionDuy Nguyen, Archiki Prasad, Elias Stengel-Eskin, Mohit BansalACL 2025 · 被引用 30 次
- Learning to Steer: Input-dependent Steering for Multimodal LLMsJayneel Parekh, Pegah Khayatan, Mustafa Shukor, Arnaud Dapogny 等NeurIPS 2025 · 被引用 14 次
