Learning Multi-Level Features with Matryoshka Sparse Autoencoders
Bart Bussmann, Noa Nabeshima, Adam Karvonen, Neel Nanda
摘要
Sparse autoencoders (SAEs) have emerged as a powerful tool for interpreting neural networks by extracting the concepts represented in their activations. However, choosing the size of the SAE dictionary (i.e. number of learned concepts) creates a tension: as dictionary size increases to capture more relevant concepts, sparsity incentivizes features to be split or absorbed into more specific features, leaving high-level features missing or warped. We introduce Matryoshka SAEs, a novel variant that addresses these issues by simultaneously training multiple nested dictionaries of increasing size, forcing the smaller dictionaries to independently reconstruct the inputs without using the larger dictionaries. This organizes features hierarchically -the smaller dictionaries learn general concepts, while the larger dictionaries learn more specific concepts, without incentive to absorb the high-level features. We train Matryoshka SAEs on Gemma-2-2B and TinyStories and find superior performance on sparse probing and targeted concept erasure tasks, more disentangled concept representations, and reduced feature absorption. While there is a minor tradeoff with reconstruction performance, we believe Matryoshka SAEs are a superior alternative for practical tasks, as they enable training arbitrarily large SAEs while retaining interpretable features at different levels of abstraction.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper33
- From Flat to Hierarchical: Extracting Sparse Representations with Matching PursuitValérie Costa, Thomas Fel, Ekdeep Singh Lubana, Bahareh Tolooshams 等NeurIPS 2025 · 被引用 54 次
- What's In My Human Feedback? Learning Interpretable Descriptions of Preference DataRajiv Movva, Smitha Milli, Sewon Min, Emma PiersonICLR 2026 · 被引用 27 次
- Priors in time: Missing inductive biases for language model interpretabilityEkdeep Singh Lubana, Can Rager, Sai Sumedh R. Hindupur, Valérie Costa 等ICLR 2026 · 被引用 19 次
- Dense SAE Latents Are Features, Not BugsXiaoqing Sun, Alessandro Stolfo, Joshua Engels, Ben Wu 等NeurIPS 2025 · 被引用 19 次
- Temporal Sparse Autoencoders: Leveraging the Sequential Nature of Language for InterpretabilityUsha Bhalla, Alex Oesterling, Claudio Mayrink Verdun, Himabindu Lakkaraju 等ICLR 2026 · 被引用 18 次
它引用的顶会 Paper8
- Sparse Autoencoders Find Highly Interpretable Features in Language ModelsRobert Huben, Hoagy Cunningham, Logan Riggs Smith, Aidan Ewart 等ICLR 2024 · 被引用 1,072 次
- Matryoshka Representation LearningAditya Kusupati, Gantavya Bhatt, Aniket Rege, Matthew Wallingford 等NeurIPS 2022 · 被引用 364 次
- A is for Absorption: Studying Feature Splitting and Absorption in Sparse AutoencodersDavid Chanin, James Wilken-Smith, Tomás Dulka, Hardik Bhatnagar 等NeurIPS 2025 · 被引用 168 次
- Measuring Progress in Dictionary Learning for Language Model Interpretability with Board Game ModelsAdam Karvonen, Benjamin Wright, Can Rager, Rico Angell 等NeurIPS 2024 · 被引用 66 次
- Scaling and evaluating sparse autoencodersLeo Gao, Tom Dupré la Tour, Henk Tillman, Gabriel Goh 等ICLR 2025 · 被引用 10 次
相关 Paper
- Interpreting CLIP with Hierarchical Sparse AutoencodersVladimir Zaigrajew, Hubert Baniecki, Przemyslaw BiecekICML 2025
- Projecting Assumptions: The Duality Between Sparse Autoencoders and Concept GeometrySai Sumedh R. Hindupur, Ekdeep Singh Lubana, Thomas Fel, Demba BaNeurIPS 2025 · 被引用 65 次
- Toward Identifiable Sparse AutoencodersWalter Nelson, Theofanis Karaletsos, Francesco LocatelloICML 2026 · 被引用 1 次
- SAEBench: A Comprehensive Benchmark for Sparse Autoencoders in Language Model InterpretabilityAdam Karvonen, Can Rager, Johnny Lin, Curt Tigges 等ICML 2025
- Ensembling Sparse AutoencodersSoham Gadgil, Chris Lin, Su-In LeeICML 2026
