CR: Cross-sample Consistency Regularization Mitigates Feature Splitting and Absorption in Sparse Autoencoders
Haoran Jin, Xiting Wang, Shijie Ren, Hong Xie, Defu Lian
Abstract
Sparse Autoencoders (SAEs) are widely used to interpret large language models by decomposing activations into sparse, human-understandable features, but scaling to large dictionaries exposes fundamental challenges. Systematic studies reveal pervasive feature splitting that fragments coherent concepts into non-atomic latents and widespread feature absorption that creates arbitrary exceptions in general features, severely compromising latent reliability. These issues stem from inconsistent latent assignment across samples: without cross-sample constraints, per-sample optimization often allows a single underlying concept to be inconsistently distributed across multiple redundant or interfering latents. To address this, we introduce CR (Cross-sample Consistency Regularization). CR explicitly encourages that each semantic feature is consistently represented by a unified latent across the batch by penalizing the co-activation of directionally similar latents. Comprehensive evaluation demonstrates that CR effectively mitigates both splitting and absorption while, crucially, preserving reconstruction fidelity, providing a principled solution that enhances latent interpretability without degrading model performance. Source code is available.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 5ee9f279-afe7-4e52-ba33-cf0e1734782eBuilds on9
- Sparse Autoencoders Find Highly Interpretable Features in Language ModelsRobert Huben, Hoagy Cunningham, Logan Riggs Smith, Aidan Ewart et al.ICLR 2024 · 1,072 citations
- A is for Absorption: Studying Feature Splitting and Absorption in Sparse AutoencodersDavid Chanin, James Wilken-Smith, Tomás Dulka, Hardik Bhatnagar et al.NeurIPS 2025 · 168 citations
- Scaling and evaluating sparse autoencodersLeo Gao, Tom Dupré la Tour, Henk Tillman, Gabriel Goh et al.ICLR 2025 · 10 citations
- Do I Know This Entity? Knowledge Awareness and Hallucinations in Language ModelsJavier Ferrando, Oscar Balcells Obeso, Senthooran Rajamanoharan, Neel NandaICLR 2025 · 1 citation
- Automatically Interpreting Millions of Features in Large Language ModelsGonçalo Paulo, Alex Mallen, Caden Juang, Nora BelroseICML 2025
Related papers
- Mechanistic Interpretability Should Prioritize Feature Consistency in Sparse AutoencodersXiangchen Song, Aashiq Muhamed, Yujia Zheng, Lingjing Kong et al.ACL 2026
- AbsTopK: Rethinking Sparse Autoencoders For Bidirectional FeaturesXudong Zhu, Mohammad Mahdi Khalili, Zhihui ZhuICLR 2026 · 10 citations
- Temporal Sparse Autoencoders: Leveraging the Sequential Nature of Language for InterpretabilityUsha Bhalla, Alex Oesterling, Claudio Mayrink Verdun, Himabindu Lakkaraju et al.ICLR 2026 · 18 citations
- Revising and Falsifying Sparse Autoencoder Feature ExplanationsGeorge Ma, Samuel Pfrommer, Somayeh SojoudiNeurIPS 2025 · 7 citations
- Archetypal SAE: Adaptive and Stable Dictionary Learning for Concept Extraction in Large Vision ModelsThomas Fel, Ekdeep Singh Lubana, Jacob S. Prince, Matthew Kowal et al.ICML 2025
