AbsTopK: Rethinking Sparse Autoencoders For Bidirectional Features
Xudong Zhu, Mohammad Mahdi Khalili, Zhihui Zhu
Abstract
Sparse autoencoders (SAEs) have emerged as powerful techniques for interpretability of large language models (LLMs), aiming to decompose hidden states into meaningful semantic features. While several SAE variants have been proposed, there remains no principled framework to derive SAEs from the original dictionary learning formulation. In this work, we introduce such a framework by unrolling the proximal gradient method for sparse coding. We show that a single-step update naturally recovers common SAE variants, including ReLU, JumpReLU, and TopK. Through this lens, we reveal a fundamental limitation of existing SAEs: their sparsity-inducing regularizers enforce non-negativity, preventing a single feature from representing bidirectional concepts (e.g., male vs. female). This structural constraint fragments semantic axes into separate, redundant features, limiting representational completeness. To address this issue, we propose AbsTopK SAE, a new variant derived from the sparsity constraint that applies hard thresholding over the largest-magnitude activations. By preserving both positive and negative activations, AbsTopK uncovers richer, bidirectional conceptual representations. Comprehensive experiments across multiple LLMs and seven probing and steering tasks show that AbsTopK improves reconstruction fidelity, enhances interpretability, and enables single features to encode contrasting concepts. Remarkably, AbsTopK matches or even surpasses the Difference-in-Mean method—a supervised approach that requires labeled data for each concept and has been shown in prior work to outperform SAEs.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 6ecedbb9-31a0-4ddc-a2cf-4621484705bcCited by top-tier papers2
- Toward Identifiable Sparse AutoencodersWalter Nelson, Theofanis Karaletsos, Francesco LocatelloICML 2026 · 1 citation
- Mechanistic Interpretability Should Prioritize Feature Consistency in Sparse AutoencodersXiangchen Song, Aashiq Muhamed, Yujia Zheng, Lingjing Kong et al.ACL 2026
Builds on25
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou et al.ICLR 2021 · 7,905 citations
- Pythia: A Suite for Analyzing Large Language Models Across Training and ScalingStella Biderman, Hailey Schoelkopf, Quentin Gregory Anthony, Herbie Bradley et al.ICML 2023 · 1,822 citations
- Refusal in Language Models Is Mediated by a Single DirectionAndy Arditi, Oscar Obeso, Aaquib Syed, Daniel Paleka et al.NeurIPS 2024 · 1,166 citations
- Sparse Autoencoders Find Highly Interpretable Features in Language ModelsRobert Huben, Hoagy Cunningham, Logan Riggs Smith, Aidan Ewart et al.ICLR 2024 · 1,072 citations
- HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust RefusalMantas Mazeika, Long Phan, Xuwang Yin, Andy Zou et al.ICML 2024 · 1,031 citations
Related papers
- Temporal Sparse Autoencoders: Leveraging the Sequential Nature of Language for InterpretabilityUsha Bhalla, Alex Oesterling, Claudio Mayrink Verdun, Himabindu Lakkaraju et al.ICLR 2026 · 18 citations
- Step-Level Sparse Autoencoder for Reasoning Process InterpretationXuan Yang, Jiayu Liu, Yuhang Lai, Hao Xu et al.ICML 2026 · 2 citations
- Sparse but Wrong: Incorrect L0 Leads to Incorrect Features in Sparse AutoencodersDavid Chanin, Adrià Garriga-AlonsoICML 2026 · 8 citations
- CR: Cross-sample Consistency Regularization Mitigates Feature Splitting and Absorption in Sparse AutoencodersHaoran Jin, Xiting Wang, Shijie Ren, Hong Xie et al.ICML 2026
- DLM-Scope: Mechanistic Interpretability of Diffusion Language Models via Sparse AutoencodersXu Wang, Bingqing Jiang, Yu Wan, Baosong Yang et al.ICML 2026
