A Geometric Information Bottleneck for Activation Steering
Toan Doan, Thin Nguyen, Sunil Gupta
摘要
Activation-based steering methods for large language models often induce broad, entangled changes in model behavior, inadvertently altering capabilities unrelated to the intended behavior, which limits their reliability for fine-grained behavioral control. We address this limitation by reframing behavioral intervention through a geometric information bottleneck (IB) perspective, in which effective steering corresponds to selectively modifying task-relevant information while preserving the geometric structure of orthogonal representational subspaces. Building on this view, we propose a disentanglement-based intervention framework, termed IB-ACT, that identifies both where and how to intervene by exploiting the layer-wise geometry of representation spaces to isolate behavior-relevant information without disturbing other dimensions. Our method introduces a layer-selection mechanism determined prior to intervention, rather than relying on post hoc sparsity or regularization losses, and applies geometrically constrained transformations that target behavior-relevant subspaces in activation space while preserving orthogonal structure. We provide theoretical justification showing that interventions at these layers reduce unintended information leakage under an IB-style objective. Empirically, we evaluate IB-ACT on toxicity control and hallucination reduction in large language models and demonstrate consistent improvements over recent baselines, while analyzing the spectral structure of behavior-relevant representations for jailbreak mitigation. Overall, our findings suggest that selectively intervening at structurally appropriate layers is critical for controllable and disentangled behavioral steering in large language models.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
相关 Paper
- Breaking Bad Tokens: Detoxification of LLMs Using Sparse AutoencodersAgam Goyal, Vedant Rathi, William Yeh, Yian Wang 等EMNLP 2025 · 被引用 1 次
- Jailbreaking the Matrix: Nullspace Steering for Controlled Model SubversionVishal Pramanik, Maisha Maliha, Susmit Jha, Sumit Kumar JhaICLR 2026 · 被引用 3 次
- Redefining Experts: Interpretable Decomposition of Language Models for Toxicity MitigationZuhair Hasan Shaik, Abdullah Mazhar, Aseem Srivastava, Md. Shad AkhtarNeurIPS 2025 · 被引用 3 次
- Detecting What Queries Seek: Steering LLM Safety with FFN Output Activation MonitoringXiaohao Luo, Ying Wei, Rui ZhaoACL 2026
- Controlling Language and Diffusion Models by Transporting ActivationsPau Rodríguez, Arno Blaas, Michal Klein, Luca Zappella 等ICLR 2025
