A Geometric Information Bottleneck for Activation Steering
Toan Doan, Thin Nguyen, Sunil Gupta
Abstract
Activation-based steering methods for large language models often induce broad, entangled changes in model behavior, inadvertently altering capabilities unrelated to the intended behavior, which limits their reliability for fine-grained behavioral control. We address this limitation by reframing behavioral intervention through a geometric information bottleneck (IB) perspective, in which effective steering corresponds to selectively modifying task-relevant information while preserving the geometric structure of orthogonal representational subspaces. Building on this view, we propose a disentanglement-based intervention framework, termed IB-ACT, that identifies both where and how to intervene by exploiting the layer-wise geometry of representation spaces to isolate behavior-relevant information without disturbing other dimensions. Our method introduces a layer-selection mechanism determined prior to intervention, rather than relying on post hoc sparsity or regularization losses, and applies geometrically constrained transformations that target behavior-relevant subspaces in activation space while preserving orthogonal structure. We provide theoretical justification showing that interventions at these layers reduce unintended information leakage under an IB-style objective. Empirically, we evaluate IB-ACT on toxicity control and hallucination reduction in large language models and demonstrate consistent improvements over recent baselines, while analyzing the spectral structure of behavior-relevant representations for jailbreak mitigation. Overall, our findings suggest that selectively intervening at structurally appropriate layers is critical for controllable and disentangled behavioral steering in large language models.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 1cf4bb3f-3784-439a-9079-057d33080300Related papers
- Breaking Bad Tokens: Detoxification of LLMs Using Sparse AutoencodersAgam Goyal, Vedant Rathi, William Yeh, Yian Wang et al.EMNLP 2025 · 1 citation
- Jailbreaking the Matrix: Nullspace Steering for Controlled Model SubversionVishal Pramanik, Maisha Maliha, Susmit Jha, Sumit Kumar JhaICLR 2026 · 3 citations
- Redefining Experts: Interpretable Decomposition of Language Models for Toxicity MitigationZuhair Hasan Shaik, Abdullah Mazhar, Aseem Srivastava, Md. Shad AkhtarNeurIPS 2025 · 3 citations
- Detecting What Queries Seek: Steering LLM Safety with FFN Output Activation MonitoringXiaohao Luo, Ying Wei, Rui ZhaoACL 2026
- Controlling Language and Diffusion Models by Transporting ActivationsPau Rodríguez, Arno Blaas, Michal Klein, Luca Zappella et al.ICLR 2025
