CR: Cross-sample Consistency Regularization Mitigates Feature Splitting and Absorption in Sparse Autoencoders
Haoran Jin, Xiting Wang, Shijie Ren, Hong Xie, Defu Lian
摘要
Sparse Autoencoders (SAEs) are widely used to interpret large language models by decomposing activations into sparse, human-understandable features, but scaling to large dictionaries exposes fundamental challenges. Systematic studies reveal pervasive feature splitting that fragments coherent concepts into non-atomic latents and widespread feature absorption that creates arbitrary exceptions in general features, severely compromising latent reliability. These issues stem from inconsistent latent assignment across samples: without cross-sample constraints, per-sample optimization often allows a single underlying concept to be inconsistently distributed across multiple redundant or interfering latents. To address this, we introduce CR (Cross-sample Consistency Regularization). CR explicitly encourages that each semantic feature is consistently represented by a unified latent across the batch by penalizing the co-activation of directionally similar latents. Comprehensive evaluation demonstrates that CR effectively mitigates both splitting and absorption while, crucially, preserving reconstruction fidelity, providing a principled solution that enhances latent interpretability without degrading model performance. Source code is available.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper9
- Sparse Autoencoders Find Highly Interpretable Features in Language ModelsRobert Huben, Hoagy Cunningham, Logan Riggs Smith, Aidan Ewart 等ICLR 2024 · 被引用 1,072 次
- A is for Absorption: Studying Feature Splitting and Absorption in Sparse AutoencodersDavid Chanin, James Wilken-Smith, Tomás Dulka, Hardik Bhatnagar 等NeurIPS 2025 · 被引用 168 次
- Scaling and evaluating sparse autoencodersLeo Gao, Tom Dupré la Tour, Henk Tillman, Gabriel Goh 等ICLR 2025 · 被引用 10 次
- Do I Know This Entity? Knowledge Awareness and Hallucinations in Language ModelsJavier Ferrando, Oscar Balcells Obeso, Senthooran Rajamanoharan, Neel NandaICLR 2025 · 被引用 1 次
- Automatically Interpreting Millions of Features in Large Language ModelsGonçalo Paulo, Alex Mallen, Caden Juang, Nora BelroseICML 2025
相关 Paper
- Mechanistic Interpretability Should Prioritize Feature Consistency in Sparse AutoencodersXiangchen Song, Aashiq Muhamed, Yujia Zheng, Lingjing Kong 等ACL 2026
- AbsTopK: Rethinking Sparse Autoencoders For Bidirectional FeaturesXudong Zhu, Mohammad Mahdi Khalili, Zhihui ZhuICLR 2026 · 被引用 10 次
- Temporal Sparse Autoencoders: Leveraging the Sequential Nature of Language for InterpretabilityUsha Bhalla, Alex Oesterling, Claudio Mayrink Verdun, Himabindu Lakkaraju 等ICLR 2026 · 被引用 18 次
- Revising and Falsifying Sparse Autoencoder Feature ExplanationsGeorge Ma, Samuel Pfrommer, Somayeh SojoudiNeurIPS 2025 · 被引用 7 次
- Archetypal SAE: Adaptive and Stable Dictionary Learning for Concept Extraction in Large Vision ModelsThomas Fel, Ekdeep Singh Lubana, Jacob S. Prince, Matthew Kowal 等ICML 2025
