Weight-sparse transformers have interpretable circuits
Leo Gao, Achyuta Rajaram, Jacob Coxon, Soham Govande, Bowen Baker, Daniel Mossing
摘要
Finding human-understandable circuits in language models is a central goal of the field of mechanistic interpretability. We train models to have more understandable circuits by constraining most of their weights to be zeros, so that each neuron only has a few connections. To recover fine-grained circuits underlying each of several hand-crafted tasks, we prune the models to isolate the part responsible for the task. These circuits often contain neurons and residual channels that correspond to natural concepts, with a small number of straightforwardly interpretable connections between them. We study how these models scale and find that making weights sparser trades off capability for interpretability, and scaling model size improves the capability-interpretability frontier. However, scaling sparse models beyond tens of millions of nonzero parameters while preserving interpretability remains a challenge. In addition to training weight-sparse models de novo, we show preliminary results suggesting our method can also be adapted to explain existing dense models. Our work produces circuits that achieve an unprecedented level of human understandability and validates them with considerable rigor.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper7
- The Expert Strikes Back: Interpreting Mixture-of-Experts Language Models at Expert LevelJeremy Herbst, Stefan Wermter, Jae Hee LeeICML 2026 · 被引用 9 次
- SafeSeek: Universal Attribution of Safety Circuits in Language ModelsMiao Yu, Siyuan Fu, Moayad Aloqaily, Zhenhong Zhou 等ICML 2026 · 被引用 3 次
- From Weights to Activations: Is Steering the Next Frontier of Adaptation?Simon Ostermann, Daniil Gurgurov, Tanja Baeumel, Michael A. Hedderich 等ACL 2026 · 被引用 3 次
- Towards Intrinsic Interpretability of Large Language Models: A Survey of Design Principles and ArchitecturesYutong Gao, Qinglin Meng, Yuan Zhou, Liangming PanACL 2026 · 被引用 3 次
- Language Model Circuits Are Sparse in the Neuron BasisAryaman Arora, Zhengxuan Wu, Jacob Steinhardt, Sarah SchwettmannICML 2026
它引用的顶会 Paper19
- Efficient Streaming Language Models with Attention SinksGuangxuan Xiao, Yuandong Tian, Beidi Chen, Song Han 等ICLR 2024 · 被引用 1,714 次
- Towards Automated Circuit Discovery for Mechanistic InterpretabilityArthur Conmy, Augustine N. Mavor-Parker, Aengus Lynch, Stefan Heimersheim 等NeurIPS 2023 · 被引用 861 次
- A Simple and Effective Pruning Approach for Large Language ModelsMingjie Sun, Zhuang Liu, Anna Bair, J. Zico KolterICLR 2024 · 被引用 794 次
- Rigging the Lottery: Making All Tickets WinnersUtku Evci, Trevor Gale, Jacob Menick, Pablo Samuel Castro 等ICML 2020 · 被引用 723 次
- Optimal Brain Compression: A Framework for Accurate Post-Training Quantization and PruningElias Frantar, Dan AlistarhNeurIPS 2022 · 被引用 440 次
相关 Paper
- Sparse Feature Circuits: Discovering and Editing Interpretable Causal Graphs in Language ModelsSamuel Marks, Can Rager, Eric J. Michaud, Yonatan Belinkov 等ICLR 2025
- Extracting Interpretable Task-Specific Circuits from Large Language Models for Faster InferenceJorge García-Carrasco, Alejandro Maté, Juan TrujilloAAAI 2025 · 被引用 3 次
- Transcoders find interpretable LLM feature circuitsJacob Dunefsky, Philippe Chlenski, Neel NandaNeurIPS 2024 · 被引用 222 次
- Finding Transformer Circuits With Edge PruningAdithya Bhaskar, Alexander Wettig, Dan Friedman, Danqi ChenNeurIPS 2024 · 被引用 72 次
- Decomposing Representation Space into Interpretable Subspaces with Unsupervised LearningXinting Huang, Michael HahnICLR 2026 · 被引用 7 次
