Finding Transformer Circuits With Edge Pruning
Adithya Bhaskar, Alexander Wettig, Dan Friedman, Danqi Chen
摘要
The path to interpreting a language model often proceeds via analysis of circuits -- sparse computational subgraphs of the model that capture specific aspects of its behavior. Recent work has automated the task of discovering circuits. Yet, these methods have practical limitations, as they rely either on inefficient search algorithms or inaccurate approximations. In this paper, we frame automated circuit discovery as an optimization problem and propose Edge Pruning as an effective and scalable solution. Edge Pruning leverages gradient-based pruning techniques, but instead of removing neurons or components, it prunes the edges between components. Our method finds circuits in GPT-2 that use less than half the number of edges compared to circuits found by previous methods while being equally faithful to the full model predictions on standard circuit-finding tasks. Edge Pruning is efficient even with as many as 100K examples, outperforming previous methods in speed and producing substantially better circuits. It also perfectly recovers the ground-truth circuits in two models compiled with Tracr. Thanks to its efficiency, we scale Edge Pruning to CodeLlama-13B, a model over 100x the scale that prior methods operate on. We use this setting for a case study comparing the mechanisms behind instruction prompting and in-context learning. We find two circuits with more than 99.96% sparsity that match the performance of the full model and reveal that the mechanisms in the two settings overlap substantially. Our case study shows that Edge Pruning is a practical and scalable tool for interpretability and sheds light on behaviors that only emerge in large models.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper30
- Knowledge Circuits in Pretrained TransformersYunzhi Yao, Ningyu Zhang, Zekun Xi, Mengru Wang 等NeurIPS 2024 · 被引用 71 次
- Optimal ablation for interpretabilityMaximilian Li, Lucas JansonNeurIPS 2024 · 被引用 32 次
- Mechanistic Interpretability as Statistical Estimation: A Variance AnalysisMaxime Méloux, François Portet, Maxime PeyrardICML 2026 · 被引用 13 次
- Sheaf Discovery with Joint Computation Graph Pruning and Flexible GranularityLei Yu, Jingcheng Niu, Zining Zhu, Xi Chen 等EMNLP 2025 · 被引用 11 次
- Rethinking Circuit Completeness in Language Models: AND, OR, and ADDER GatesHang Chen, Jiaying Zhu, Xinyu Yang, Wenya WangNeurIPS 2025 · 被引用 11 次
它引用的顶会 Paper19
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Sparse Autoencoders Find Highly Interpretable Features in Language ModelsRobert Huben, Hoagy Cunningham, Logan Riggs Smith, Aidan Ewart 等ICLR 2024 · 被引用 1,072 次
- Towards Automated Circuit Discovery for Mechanistic InterpretabilityArthur Conmy, Augustine N. Mavor-Parker, Aengus Lynch, Stefan Heimersheim 等NeurIPS 2023 · 被引用 861 次
- Investigating Gender Bias in Language Models Using Causal Mediation AnalysisJesse Vig, Sebastian Gehrmann, Yonatan Belinkov, Sharon Qian 等NeurIPS 2020 · 被引用 851 次
- Sheared LLaMA: Accelerating Language Model Pre-training via Structured PruningMengzhou Xia, Tianyu Gao, Zhiyuan Zeng, Danqi ChenICLR 2024 · 被引用 453 次
相关 Paper
- Weight-sparse transformers have interpretable circuitsLeo Gao, Achyuta Rajaram, Jacob Coxon, Soham Govande 等ICML 2026
- Sparse Feature Circuits: Discovering and Editing Interpretable Causal Graphs in Language ModelsSamuel Marks, Can Rager, Eric J. Michaud, Yonatan Belinkov 等ICLR 2025
- Efficient Automated Circuit Discovery in Transformers using Contextual DecompositionAliyah R. Hsu, Georgia Zhou, Yeshwanth Cherapanamjeri, Yaxuan Huang 等ICLR 2025
- EfficientLLM: Unified Pruning-Aware Pretraining for Auto-Designed Compact Language ModelsXingrun Xing, Zheng Liu, Shitao Xiao, Boyan Gao 等ACL 2026
- Scaling Sparse Feature Circuits For Studying In-Context LearningDmitrii Kharlapenko, Stepan Shabalin, Arthur Conmy, Neel NandaICML 2025
