Compute Optimal Inference and Provable Amortisation Gap in Sparse Autoencoders
Charles O'Neill, Alim Gumran, David A. Klindt
摘要
A recent line of work has shown promise in using sparse autoencoders (SAEs) to uncover interpretable features in neural network representations. However, the simple linear-nonlinear encoding mechanism in SAEs limits their ability to perform accurate sparse inference. Using compressed sensing theory, we prove that an SAE encoder is inherently insufficient for accurate sparse inference, even in solvable cases. We then decouple encoding and decoding processes to empirically explore conditions where more sophisticated sparse inference methods outperform traditional SAE encoders. Our results reveal substantial performance gains with minimal compute increases in correct inference of sparse codes. We demonstrate this generalises to SAEs applied to large language models, where more expressive encoders achieve greater interpretability. This work opens new avenues for understanding neural network representations and analysing large language model activations. Recent work has investigated the "superposition hypothe- Compute Optimal Inference and Provable Amortisation Gap in Sparse Autoencoders
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- AbsTopK: Rethinking Sparse Autoencoders For Bidirectional FeaturesXudong Zhu, Mohammad Mahdi Khalili, Zhihui ZhuICLR 2026 · 被引用 10 次
- Brain-like Variational InferenceHadi Vafaii, Dekel Galor, Jacob L. YatesNeurIPS 2025 · 被引用 7 次
- The Price of Amortized inference in Sparse AutoencodersWenjie Sun, Di Wang, Lijie HuICLR 2026
- Towards Atoms of Large Language ModelsChenhui Hu, Pengfei Cao, Yubo Chen, Kang Liu 等ICML 2026
- Mechanistic Interpretability Should Prioritize Feature Consistency in Sparse AutoencodersXiangchen Song, Aashiq Muhamed, Yujia Zheng, Lingjing Kong 等ACL 2026
它引用的顶会 Paper4
- Sparse Autoencoders Find Highly Interpretable Features in Language ModelsRobert Huben, Hoagy Cunningham, Logan Riggs Smith, Aidan Ewart 等ICLR 2024 · 被引用 1,072 次
- Scaling and evaluating sparse autoencodersLeo Gao, Tom Dupré la Tour, Henk Tillman, Gabriel Goh 等ICLR 2025 · 被引用 10 次
- The Geometry of Categorical and Hierarchical Concepts in Large Language ModelsKiho Park, Yo Joong Choe, Yibo Jiang, Victor VeitchICLR 2025 · 被引用 3 次
- Towards Principled Evaluations of Sparse Autoencoders for Interpretability and ControlAleksandar Makelov, Georg Lange, Neel NandaICLR 2025
相关 Paper
- On the Limits of Sparse Autoencoders: A Theoretical Framework and Reweighted RemedyJingyi Cui, Qi Zhang, Yifei Wang, Yisen WangICLR 2026 · 被引用 16 次
- Improving Sparse Decomposition of Language Model Activations with Gated Sparse AutoencodersSenthooran Rajamanoharan, Arthur Conmy, Lewis Smith, Tom Lieberum 等NeurIPS 2024 · 被引用 49 次
- Are Sparse Autoencoders Useful? A Case Study in Sparse ProbingSubhash Kantamneni, Joshua Engels, Senthooran Rajamanoharan, Max Tegmark 等ICML 2025
- From Flat to Hierarchical: Extracting Sparse Representations with Matching PursuitValérie Costa, Thomas Fel, Ekdeep Singh Lubana, Bahareh Tolooshams 等NeurIPS 2025 · 被引用 54 次
- Projecting Assumptions: The Duality Between Sparse Autoencoders and Concept GeometrySai Sumedh R. Hindupur, Ekdeep Singh Lubana, Thomas Fel, Demba BaNeurIPS 2025 · 被引用 65 次
