Beyond Interpretability: The Gains of Feature Monosemanticity on Model Robustness
Qi Zhang, Yifei Wang, Jingyi Cui, Xiang Pan, Qi Lei, Stefanie Jegelka, Yisen Wang
摘要
Deep learning models often suffer from a lack of interpretability due to polysemanticity, where individual neurons are activated by multiple unrelated semantics, resulting in unclear attributions of model behavior. Recent advances in monosemanticity, where neurons correspond to consistent and distinct semantics, have significantly improved interpretability but are commonly believed to compromise accuracy. In this work, we challenge the prevailing belief of the accuracy-interpretability tradeoff, showing that monosemantic features not only enhance interpretability but also bring concrete gains in model performance. Across multiple robust learning scenarios-including input and label noise, fewshot learning, and out-of-domain generalization-our results show that models leveraging monosemantic features significantly outperform those relying on polysemantic features. Furthermore, we provide empirical and theoretical understandings on the robustness gains of feature monosemanticity. Our preliminary analysis suggests that monosemanticity, by promoting better separation of feature representations, leads to more robust decision boundaries. This diverse evidence highlights the generality of monosemanticity in improving model robustness. As a first step in this new direction, we embark on exploring the learning benefits of monosemanticity beyond interpretability, supporting the longstanding hypothesis of linking interpretability and robustness. Code is available at https://github.com/PKU-ML/Beyond_Interpretability .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper7
- On the Limits of Sparse Autoencoders: A Theoretical Framework and Reweighted RemedyJingyi Cui, Qi Zhang, Yifei Wang, Yisen WangICLR 2026 · 被引用 16 次
- Finding and Reactivating Post-Trained LLMs' Hidden Safety MechanismsMingjie Li, Wai Man Si, Michael Backes, Yang Zhang 等NeurIPS 2025 · 被引用 4 次
- Visual Self-Fulfilling Alignment: Shaping Safety-Oriented Personas via Threat-Related ImagesQishun Yang, Shu Yang, Lijie Hu, Di WangACL 2026 · 被引用 1 次
- Extracting Interaction-Aware Monosemantic Concepts in Recommender SystemsDor Arviv, Yehonatan Elisha, Oren Barkan, Noam KoenigsteinAAAI 2026
- PerFit: Exploring Personalization Shifts in Representation Space of LLMsJiahong Liu, Wenhao Yu, Quanyu Dai, Zhongyang Li 等ICLR 2026
它引用的顶会 Paper10
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 被引用 24,064 次
- Symmetric Cross Entropy for Robust Learning With Noisy LabelsYisen Wang, Xingjun Ma, Zaiyi Chen, Yuan Luo 等ICCV 2019 · 被引用 1,125 次
- Sparse Autoencoders Find Highly Interpretable Features in Language ModelsRobert Huben, Hoagy Cunningham, Logan Riggs Smith, Aidan Ewart 等ICLR 2024 · 被引用 1,072 次
- Normalized Loss Functions for Deep Learning with Noisy LabelsXingjun Ma, Hanxun Huang, Yisen Wang, Simone Romano 等ICML 2020 · 被引用 547 次
- A Meta-Transfer Objective for Learning to Disentangle Causal MechanismsYoshua Bengio, Tristan Deleu, Nasim Rahaman, Nan Rosemary Ke 等ICLR 2020 · 被引用 371 次
相关 Paper
- Capturing Polysemanticity with PRISM: A Multi-Concept Feature Description FrameworkLaura Kopf, Nils Feldhus, Kirill Bykov, Philine Lou Bommer 等NeurIPS 2025 · 被引用 12 次
- Rethinking Evaluation of Sparse Autoencoders through the Representation of Polysemous WordsGouki Minegishi, Hiroki Furuta, Yusuke Iwasawa, Yutaka MatsuoICLR 2025
- Robust Models Are More Interpretable Because Attributions Look NormalZifan Wang, Matt Fredrikson, Anupam DattaICML 2022 · 被引用 33 次
- LACMUS: Latent Concept Masking for General Robustness Enhancement of DNNsShuo Wang, Hongsheng Hu, Jiamin Chang, Benjamin Zi Hao Zhao 等S&P 2024
- Connecting Interpretability and Robustness in Decision Trees through SeparationMichal Moshkovitz, Yao-Yuan Yang, Kamalika ChaudhuriICML 2021 · 被引用 28 次
