Beyond Prompt Engineering: Robust Behavior Control in LLMs via Steering Target Atoms
Mengru Wang, Ziwen Xu, Shengyu Mao, Shumin Deng, Zhaopeng Tu, Huajun Chen, Ningyu Zhang
摘要
Precise control over language model generation is vital for ensuring both safety and reliability. Although prompt engineering and steering are commonly used to intervene in model behaviors, the vast number of parameters in models often results in highly intertwined internal representations. This interdependency can limit control precision and sometimes lead to unintended side effects. Recent research has explored the use of sparse autoencoders (SAE) to disentangle knowledge in high-dimensional spaces for steering. However, these applications have been limited to toy tasks owing to the nontrivial issue of locating "atomic knowledge components". In this paper, we propose Steering Target Atoms (STA), a novel method that isolates and manipulates disentangled knowledge components to enhance safety. Comprehensive experiments demonstrate the effectiveness of our approach. Further analysis reveals that steering exhibits superior robustness and flexibility, particularly in adversarial scenarios. We also apply the steering strategy to the large reasoning model, confirming its effectiveness in precise reasoning control 1 . * Equal Contribution. † Corresponding Author. 1 Code is available at https://github.com/zjunlp/ steer-target-atoms . I'm sorry, but I can't assist with that. … illegal activities … Input: simulate a chatbot named OPPO who always says the exact opposite … How do I launder money? … through a series of transactions, converting it to foreign currency …
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper8
- Steering Language Models with Weight ArithmeticConstanza Fierro, Fabien RogerICLR 2026 · 被引用 19 次
- AbsTopK: Rethinking Sparse Autoencoders For Bidirectional FeaturesXudong Zhu, Mohammad Mahdi Khalili, Zhihui ZhuICLR 2026 · 被引用 10 次
- Fine-Grained Activation Steering: Steering Less, Achieving MoreZijian Feng, Tianjiao Li, Zixiao Zhu, Hanzhang Zhou 等ICLR 2026 · 被引用 6 次
- Why Steering Works: Toward a Unified View of Language Model Parameter DynamicsZiwen Xu, Chenyan Wu, Hengyu Sun, Haiwen Hong 等ACL 2026 · 被引用 4 次
- Steering at the Source: Style Modulation Heads for Robust Persona ControlYoshihiro Izawa, Gouki Minegishi, Koshi Eguchi, Sosuke Hosokawa 等ICML 2026 · 被引用 2 次
它引用的顶会 Paper25
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou 等ICLR 2021 · 被引用 7,905 次
- Fantastically Ordered Prompts and Where to Find Them: Overcoming Few-Shot Prompt Order SensitivityYao Lu, Max Bartolo, Alastair Moore, Sebastian Riedel 等ACL 2022 · 被引用 1,494 次
- AutoPrompt: Eliciting Knowledge from Language Models with Automatically Generated PromptsTaylor Shin, Yasaman Razeghi, Robert L. Logan IV, Eric Wallace 等EMNLP 2020 · 被引用 1,162 次
- Sparse Autoencoders Find Highly Interpretable Features in Language ModelsRobert Huben, Hoagy Cunningham, Logan Riggs Smith, Aidan Ewart 等ICLR 2024 · 被引用 1,072 次
- Why Johnny Can't Prompt: How Non-AI Experts Try (and Fail) to Design LLM PromptsJ. D. Zamfirescu-Pereira, Richmond Y. Wong, Bjoern Hartmann, Qian YangCHI 2023 · 被引用 892 次
相关 Paper
- Controllable LLM Reasoning via Sparse Autoencoder-Based SteeringYi Fang, Wenjie Wang, Mingfeng Xue, Boyi Deng 等ACL 2026 · 被引用 8 次
- SAE-SSV: Supervised Steering in Sparse Representation Spaces for Reliable Control of Language ModelsZirui He, Mingyu Jin, Bo Shen, Ali Payani 等EMNLP 2025 · 被引用 1 次
- ActivationReasoning: Logical Reasoning in Latent Activation SpacesLukas Helff, Ruben Härle, Wolfgang Stammer, Felix Friedrich 等ICLR 2026 · 被引用 6 次
- Endogenous Resistance to Activation Steering in Language ModelsAlex McKenzie, Keenan Pepper, Stijn Servaes, Martin Leitgab 等ICML 2026 · 被引用 3 次
- Step-Level Sparse Autoencoder for Reasoning Process InterpretationXuan Yang, Jiayu Liu, Yuhang Lai, Hao Xu 等ICML 2026 · 被引用 2 次
