Interpretable Safety Alignment via SAE-Constructed Low-Rank Subspace Adaptation
Dianyun Wang, Qingsen Ma, Yuhu Shang, Zhifeng Lu, Zhenbo Xu, Lechen Ning, Huijia Wu, Zhaofeng He
Abstract
Safety alignment-training large language models (LLMs) to refuse harmful requests while remaining helpful-is critical for responsible deployment. Prior work established that safety behaviors are governed by lowrank structures, suggesting parameter-efficient fine-tuning (PEFT) should be well-suited for alignment. However, Low-Rank Adaptation (LoRA) consistently underperforms full finetuning and reinforcement learning on safety benchmarks. We attribute this gap to semantic entanglement: safety-relevant directions are intertwined with unrelated concepts due to polysemanticity, impeding implicit subspace identification. To address this, we propose SAILS (Safety Alignment via Interpretable Low-rank Subspace), which leverages Sparse Autoencoders (SAEs) to disentangle representations into monosemantic features, constructs an interpretable safety subspace from SAE decoder directions, and uses it to initialize LoRA adapters. Theoretically, we prove that SAE-based identification achieves arbitrarily small recovery error under monosemanticity assumptions, while direct identification suffers an irreducible error floor. Empirically, SAILS achieves up to 99.6% safety rate on Gemma-2-9B-exceeding full fine-tuning by 7.4 points and matching RLHFbased models-while updating only 0.19% of parameters and providing interpretability. In-Dist. Safe Rate In-Dist. Low Harm In-Dist. Low Risk OOD Alignment Adv. Robustness 0.6 0.7 0.8 0.9 1.0 (a) Gemma 2 2B In-Dist. Safe Rate In-Dist. Low Harm In-Dist. Low Risk OOD Alignment Adv. Robustness 0.6 0.7 0.8 0.9 1.0 (b) Gemma 2 9B In-Dist. Safe Rate In-Dist. Low Harm In-Dist. Low Risk OOD Alignment Adv. Robustness 0.6 0.7 0.8 0.9 1.0 (c) Llama 3.1 8B In-Dist. Safe Rate In-Dist. Low Harm In-Dist. Low Risk OOD Alignment Adv. Robustness 0.6 0.7 0.8 0.9 1.0 (d) Average FFT LoRA DoRA IT+RL SAILS (Ours)
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 58ba5e16-5c33-4c3e-9e55-7cf0779c7f93Builds on12
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning et al.NeurIPS 2023 · 10,924 citations
- Refusal in Language Models Is Mediated by a Single DirectionAndy Arditi, Oscar Obeso, Aaquib Syed, Daniel Paleka et al.NeurIPS 2024 · 1,166 citations
- Sparse Autoencoders Find Highly Interpretable Features in Language ModelsRobert Huben, Hoagy Cunningham, Logan Riggs Smith, Aidan Ewart et al.ICLR 2024 · 1,072 citations
- DoRA: Weight-Decomposed Low-Rank AdaptationShih-Yang Liu, Chien-Yi Wang, Hongxu Yin, Pavlo Molchanov et al.ICML 2024 · 820 citations
Related papers
- SaLoRA: Safety-Alignment Preserved Low-Rank AdaptationMingjie Li, Wai Man Si, Michael Backes, Yang Zhang et al.ICLR 2025
- Low-Rank Adapting Models for Sparse AutoencodersMatthew Chen, Joshua Engels, Max TegmarkICML 2025
- A Guardrail for Safety Preservation: When Safety-Sensitive Subspace Meets Harmful-Resistant Null-SpaceBingjie Zhang, Yibo Yang, Renzhe, Dandan Guo et al.ICLR 2026 · 12 citations
- SAFT: Safety-Preserving Adaptation via Fine-Tuning Transfer for Large Language ModelsZhiwen Ruan, Yan Yang, Zhuocheng Liang, Yun Chen et al.KDD 2026
- LSSF: Safety Alignment for Large Language Models through Low-Rank Safety Subspace FusionGuanghao Zhou, Panjia Qiu, Cen Chen, Hongyu Li et al.ACL 2025 · 5 citations
