Low-Rank Adapting Models for Sparse Autoencoders
Matthew Chen, Joshua Engels, Max Tegmark
Abstract
Sparse autoencoders (SAEs) decompose language model representations into a sparse set of linear latent vectors. Recent works have improved SAEs using language model gradients, but these techniques require many expensive backward passes during training and still cause a significant increase in cross entropy loss when SAE reconstructions are inserted into the model. In this work, we improve on these limitations by taking a fundamentally different approach: we use low-rank adaptation (LoRA) to finetune the language model itself around a previously trained SAE. We analyze our method across SAE sparsity, SAE width, language model size, LoRA rank, and model layer on the Gemma Scope family of SAEs. In these settings, our method reduces the cross entropy loss gap by 30% to 55% when SAEs are inserted during the forward pass. We also find that compared to end-to-end (e2e) SAEs, our approach achieves the same downstream cross entropy loss 3× to 20× faster on Gemma-2-2B and 2× to 10× faster on Llama-3.2-1B. We further show that our technique improves downstream metrics and can adapt multiple SAEs at once without harming general language model capabilities. Our results demonstrate that improving model interpretability is not limited to post-hoc SAE training; Pareto improvements can also be achieved by directly optimizing the model itself. 1
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext ba4d5355-f172-45af-b5d9-ccdee30db581Cited by top-tier papers2
- Sparse CLIP: Co-Optimizing Interpretability and Performance in Contrastive LearningChuan Qin, Constantin Venhoff, Sonia Joseph, Fanyi Xiao et al.ICLR 2026 · 4 citations
- Breaking the Block: Preserving Data Continuity to Train Superior SAEs for Instruct ModelsJiaming Li, Haoran Ye, Yukun Chen, Xinyue Li et al.ICML 2026 · 1 citation
Builds on9
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou et al.ICLR 2021 · 7,905 citations
- Sparse Autoencoders Find Highly Interpretable Features in Language ModelsRobert Huben, Hoagy Cunningham, Logan Riggs Smith, Aidan Ewart et al.ICLR 2024 · 1,072 citations
- The Linear Representation Hypothesis and the Geometry of Large Language ModelsKiho Park, Yo Joong Choe, Victor VeitchICML 2024 · 461 citations
- A is for Absorption: Studying Feature Splitting and Absorption in Sparse AutoencodersDavid Chanin, James Wilken-Smith, Tomás Dulka, Hardik Bhatnagar et al.NeurIPS 2025 · 168 citations
Related papers
- Interpretable Safety Alignment via SAE-Constructed Low-Rank Subspace AdaptationDianyun Wang, Qingsen Ma, Yuhu Shang, Zhifeng Lu et al.ACL 2026 · 2 citations
- Dynamic Low-Rank Sparse Adaptation for Large Language ModelsWeizhong Huang, Yuxin Zhang, Xiawu Zheng, Yang Liu et al.ICLR 2025
- DenseLoRA: Dense Low-Rank Adaptation of Large Language ModelsLin Mu, Xiaoyu Wang, Li Ni, Yang Li et al.ACL 2025 · 3 citations
- SALR: Sparsity-Aware Low-Rank Representation for Efficient Fine-Tuning of Large Language ModelsLongteng Zhang, Sen Wu, Shuai Hou, Zhengyu Qing et al.AAAI 2026
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
