Skill Path: Unveiling Language Skills from Circuit Graphs
Hang Chen, Xinyu Yang, Jiaying Zhu, Wenya Wang
摘要
Circuit graph discovery has emerged as a fundamental approach to elucidating the skill mechanistic of language models. Despite the output faithfulness of circuit graphs, they suffer from atomic ablation, which causes the loss of causal dependencies between connected components. In addition, their discovery process, designed to preserve output faithfulness, inadvertently captures extraneous effects other than an isolated target skill. To alleviate these challenges, we introduce skill paths, which offers a more refined and compact representation by isolating individual skills within a linear chain of components. To enable skill path extracting from circuit graphs, we propose a three-step framework, consisting of decomposition, pruning, and post-pruning causal mediation. In particular, we offer a complete linear decomposition of the transformer model which leads to a disentangled computation graph. After pruning, we further adopt causal analysis techniques, including counterfactuals and interventions, to extract the final skill paths from the circuit graph. To underscore the significance of skill paths, we investigate three generic language skills-Previous Token Skill, Induction Skill, and In-Context Learning Skill-using our framework. Experiments support two crucial properties of these skills, namely stratification and inclusiveness. Our codes are available at: https://github.com/Zodiark-ch/Language-Skill-of-LLMs .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper10
- Towards Automated Circuit Discovery for Mechanistic InterpretabilityArthur Conmy, Augustine N. Mavor-Parker, Aengus Lynch, Stefan Heimersheim 等NeurIPS 2023 · 被引用 861 次
- Investigating Gender Bias in Language Models Using Causal Mediation AnalysisJesse Vig, Sebastian Gehrmann, Yonatan Belinkov, Sharon Qian 等NeurIPS 2020 · 被引用 851 次
- How does GPT-2 compute greater-than?: Interpreting mathematical abilities in a pre-trained language modelMichael Hanna, Ollie Liu, Alexandre VariengienNeurIPS 2023 · 被引用 251 次
- The Evolution of Statistical Induction Heads: In-Context Learning Markov ChainsEzra Edelman, Nikolaos Tsilivis, Benjamin L. Edelman, Eran Malach 等NeurIPS 2024 · 被引用 140 次
- Tracr: Compiled Transformers as a Laboratory for InterpretabilityDavid Lindner, János Kramár, Sebastian Farquhar, Matthew Rahtz 等NeurIPS 2023 · 被引用 113 次
相关 Paper
- Position-aware Automatic Circuit DiscoveryTal Haklay, Hadas Orgad, David Bau, Aaron Mueller 等ACL 2025
- Efficient Automated Circuit Discovery in Transformers using Contextual DecompositionAliyah R. Hsu, Georgia Zhou, Yeshwanth Cherapanamjeri, Yaxuan Huang 等ICLR 2025
- Towards Interpretable Sequence Continuation: Analyzing Shared Circuits in Large Language ModelsMichael Lan, Philip Torr, Fazl BarezEMNLP 2024 · 被引用 1 次
- Finding Transformer Circuits With Edge PruningAdithya Bhaskar, Alexander Wettig, Dan Friedman, Danqi ChenNeurIPS 2024 · 被引用 72 次
- EAP-GP: Mitigating Saturation Effect in Gradient-based Automated Circuit IdentificationLin Zhang, Wenshuo Dong, Zhuoran Zhang, Shu Yang 等NeurIPS 2025 · 被引用 26 次
