Toward Identifiable Sparse Autoencoders
Walter Nelson, Theofanis Karaletsos, Francesco Locatello
Abstract
Recently, sparse autoencoders (SAEs) have emerged as an attractive tool for interpreting and interacting with representations in practical neural networks. While it is common empirical folklore, we also show theoretically that SAEs are highly unstable: different training runs are likely to produce different concept dictionaries and sparse codes. We characterize the model properties that hinder the stability of real-world SAEs, and address each of these problems through minimal changes to the architecture and training procedure. Together, these changes yield two versions of an i dentifiable SAE (iSAE), a variant of the standard TopK SAE with lower reconstruction error and improved stability. We explain this improvement theoretically by connecting SAEs with traditional dictionary learning approaches, and show that the dictionaries learned in practice satisfy an approximate restricted isometry condition, rendering the corresponding sparse codes in those models near-identifiable.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 50a4ec26-710a-4b5b-a79f-ecb89fe258c6Builds on9
- Sparse Autoencoders Find Highly Interpretable Features in Language ModelsRobert Huben, Hoagy Cunningham, Logan Riggs Smith, Aidan Ewart et al.ICLR 2024 · 1,072 citations
- Sparse Autoencoders Trained on the Same Data Learn Different FeaturesGonçalo Paulo, Nora BelroseICLR 2026 · 96 citations
- Regularized Autoencoders for Isometric Representation LearningYonghyeon Lee, Sangwoong Yoon, Minjun Son, Frank Chongwoo ParkICLR 2022 · 46 citations
- AbsTopK: Rethinking Sparse Autoencoders For Bidirectional FeaturesXudong Zhu, Mohammad Mahdi Khalili, Zhihui ZhuICLR 2026 · 10 citations
- Scaling and evaluating sparse autoencodersLeo Gao, Tom Dupré la Tour, Henk Tillman, Gabriel Goh et al.ICLR 2025 · 10 citations
Related papers
- Archetypal SAE: Adaptive and Stable Dictionary Learning for Concept Extraction in Large Vision ModelsThomas Fel, Ekdeep Singh Lubana, Jacob S. Prince, Matthew Kowal et al.ICML 2025
- Projecting Assumptions: The Duality Between Sparse Autoencoders and Concept GeometrySai Sumedh R. Hindupur, Ekdeep Singh Lubana, Thomas Fel, Demba BaNeurIPS 2025 · 65 citations
- Identifying Functionally Important Features with End-to-End Sparse Dictionary LearningDan Braun, Jordan Taylor, Nicholas Goldowsky-Dill, Lee SharkeyNeurIPS 2024 · 81 citations
- Temporal Sparse Autoencoders: Leveraging the Sequential Nature of Language for InterpretabilityUsha Bhalla, Alex Oesterling, Claudio Mayrink Verdun, Himabindu Lakkaraju et al.ICLR 2026 · 18 citations
- Learning Multi-Level Features with Matryoshka Sparse AutoencodersBart Bussmann, Noa Nabeshima, Adam Karvonen, Neel NandaICML 2025
