Does Higher Interpretability Imply Better Utility? A Pairwise Analysis on Sparse Autoencoders
Xu Wang, Yan Hu, Benyou Wang, Difan Zou
摘要
Sparse Autoencoders (SAEs) are widely used to steer large language models (LLMs), based on the assumption that their interpretable features naturally enable effective model behavior steering. Yet, a fundamental question remains unanswered: does higher interpretability indeed imply better steering utility? To answer this question, we train 90 SAEs across three LLMs (Gemma-2-2B, Qwen-2.5-3B, Gemma-2-9B), spanning five architectures and six sparsity levels, and evaluate their interpretability and steering utility based on SAEBENCH [Karvonen et al., 2025] and AXBENCH [Wu et al., 2025] respectively, and perform a rankagreement analysis via Kendall's rank coefficients τ b . Based on the framework, Our analysis reveals only a relatively weak positive association (τ b ≈ 0.298), indicating that interpretability is an insufficient proxy for steering performance. We conjecture the interpretability-utility gap may stem from the selection of SAE features as not all of them are equally effective for steering. To further find features that truly steer the behavior of LLMs, we propose a novel selection criterion: ∆ Token Confidence, which measures how much amplifying a feature changes the next token distribution. We show that our method improves the steering performance of three LLMs by 52.52% compared to the current best output score-based criterion [Arad et al., 2025] . Strikingly, after selecting features with high ∆ Token Confidence, the correlation between interpretability and utility vanishes (τ b ≈ 0), and can even become negative. This further highlights the divergence between interpretability and utility for the most effective steering features.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Interpretable and Steerable Concept Bottleneck Sparse AutoencodersAkshay Kulkarni, Tsui-Wei Weng, Vivek Narayanaswamy, Shusen Liu 等CVPR 2026 · 被引用 6 次
- DLM-Scope: Mechanistic Interpretability of Diffusion Language Models via Sparse AutoencodersXu Wang, Bingqing Jiang, Yu Wan, Baosong Yang 等ICML 2026
它引用的顶会 Paper25
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- Sparse Autoencoders Find Highly Interpretable Features in Language ModelsRobert Huben, Hoagy Cunningham, Logan Riggs Smith, Aidan Ewart 等ICLR 2024 · 被引用 1,072 次
- AlpacaFarm: A Simulation Framework for Methods that Learn from Human FeedbackYann Dubois, Chen Xuechen Li, Rohan Taori, Tianyi Zhang 等NeurIPS 2023 · 被引用 948 次
- Investigating Gender Bias in Language Models Using Causal Mediation AnalysisJesse Vig, Sebastian Gehrmann, Yonatan Belinkov, Sharon Qian 等NeurIPS 2020 · 被引用 851 次
- Fantastic Generalization Measures and Where to Find ThemYiding Jiang, Behnam Neyshabur, Hossein Mobahi, Dilip Krishnan 等ICLR 2020 · 被引用 705 次
相关 Paper
- CorrSteer: Generation-Time LLM Steering via Correlated Sparse Autoencoder FeaturesSeonglae Cho, Zekun Wu, Adriano KoshiyamaICML 2026 · 被引用 5 次
- AxBench: Steering LLMs? Even Simple Baselines Outperform Sparse AutoencodersZhengxuan Wu, Aryaman Arora, Atticus Geiger, Zheng Wang 等ICML 2025
- Toward Efficient Sparse Autoencoder-Guided Steering for Improved In-Context Learning in Large Language ModelsIkhyun Cho, Julia HockenmaierEMNLP 2025 · 被引用 3 次
- Sparse Autoencoder Features for Classifications and TransferabilityJack Gallifant, Shan Chen, Kuleen Sasse, Hugo J. W. L. Aerts 等EMNLP 2025
- SAEBench: A Comprehensive Benchmark for Sparse Autoencoders in Language Model InterpretabilityAdam Karvonen, Can Rager, Johnny Lin, Curt Tigges 等ICML 2025
