A is for Absorption: Studying Feature Splitting and Absorption in Sparse Autoencoders
David Chanin, James Wilken-Smith, Tomás Dulka, Hardik Bhatnagar, Satvik Golechha, Joseph Bloom
摘要
Sparse Autoencoders (SAEs) aim to decompose the activation space of large language models (LLMs) into human-interpretable latent directions or features. As we increase the number of features in the SAE, hierarchical features tend to split into finer features ("math" may split into "algebra", "geometry", etc.), a phenomenon referred to as feature splitting. However, we show that sparse decomposition and splitting of hierarchical features is not robust. Specifically, we show that seemingly monosemantic features fail to fire where they should, and instead get "absorbed" into their children features. We coin this phenomenon feature absorption, and show that it is caused by optimizing for sparsity in SAEs whenever the underlying features form a hierarchy. We introduce a metric to detect absorption in SAEs, and validate our findings empirically on hundreds of LLM SAEs. Our investigation suggests that varying SAE sizes or sparsity is insufficient to solve this issue. We discuss the implications of feature absorption in SAEs and some potential approaches to solve the fundamental theoretical issues before SAEs can be used for interpreting LLMs robustly and at scale.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper46
- Sparse Autoencoders Trained on the Same Data Learn Different FeaturesGonçalo Paulo, Nora BelroseICLR 2026 · 被引用 96 次
- From Flat to Hierarchical: Extracting Sparse Representations with Matching PursuitValérie Costa, Thomas Fel, Ekdeep Singh Lubana, Bahareh Tolooshams 等NeurIPS 2025 · 被引用 54 次
- Detecting High-Stakes Interactions with Activation ProbesAlex McKenzie, Urja Pawar, Phil Blandfort, William Bankes 等NeurIPS 2025 · 被引用 52 次
- I Have Covered All the Bases Here: Interpreting Reasoning Features in Large Language Models via Sparse AutoencodersAndrey V. Galichin, Alexey Dontsov, Polina Druzhinina, Anton Razzhigaev 等AAAI 2026 · 被引用 31 次
- Priors in time: Missing inductive biases for language model interpretabilityEkdeep Singh Lubana, Can Rager, Sai Sumedh R. Hindupur, Valérie Costa 等ICLR 2026 · 被引用 19 次
它引用的顶会 Paper9
- Locating and Editing Factual Associations in GPTKevin Meng, David Bau, Alex Andonian, Yonatan BelinkovNeurIPS 2022 · 被引用 3,415 次
- Sparse Autoencoders Find Highly Interpretable Features in Language ModelsRobert Huben, Hoagy Cunningham, Logan Riggs Smith, Aidan Ewart 等ICLR 2024 · 被引用 1,072 次
- The Linear Representation Hypothesis and the Geometry of Large Language ModelsKiho Park, Yo Joong Choe, Victor VeitchICML 2024 · 被引用 461 次
- Measuring Progress in Dictionary Learning for Language Model Interpretability with Board Game ModelsAdam Karvonen, Benjamin Wright, Can Rager, Rico Angell 等NeurIPS 2024 · 被引用 66 次
- Emergent World Representations: Exploring a Sequence Model Trained on a Synthetic TaskKenneth Li, Aspen K. Hopkins, David Bau, Fernanda B. Viégas 等ICLR 2023 · 被引用 60 次
相关 Paper
- On the Limits of Sparse Autoencoders: A Theoretical Framework and Reweighted RemedyJingyi Cui, Qi Zhang, Yifei Wang, Yisen WangICLR 2026 · 被引用 16 次
- Rethinking Evaluation of Sparse Autoencoders through the Representation of Polysemous WordsGouki Minegishi, Hiroki Furuta, Yusuke Iwasawa, Yutaka MatsuoICLR 2025
- CR: Cross-sample Consistency Regularization Mitigates Feature Splitting and Absorption in Sparse AutoencodersHaoran Jin, Xiting Wang, Shijie Ren, Hong Xie 等ICML 2026
- Sparse but Wrong: Incorrect L0 Leads to Incorrect Features in Sparse AutoencodersDavid Chanin, Adrià Garriga-AlonsoICML 2026 · 被引用 8 次
- Sparse Autoencoders Do Not Find Canonical Units of AnalysisPatrick Leask, Bart Bussmann, Michael T. Pearce, Joseph Isaac Bloom 等ICLR 2025
