Circuit Component Reuse Across Tasks in Transformer Language Models
Jack Merullo, Carsten Eickhoff, Ellie Pavlick
摘要
Recent work in mechanistic interpretability has shown that behaviors in language models can be successfully reverse-engineered through circuit analysis. A common criticism, however, is that each circuit is task-specific, and thus such analysis cannot contribute to understanding the models at a higher level. In this work, we present evidence that insights (both low-level findings about specific heads and higher-level findings about general algorithms) can indeed generalize across tasks. Specifically, we study the circuit discovered in Wang et al. (2022) for the Indirect Object Identification (IOI) task and 1.) show that it reproduces on a larger GPT2 model, and 2.) that it is mostly reused to solve a seemingly different task: Colored Objects (Ippolito & Callison-Burch, 2023) . We provide evidence that the process underlying both tasks is functionally very similar, and contains about a 78% overlap in in-circuit attention heads. We further present a proof-of-concept intervention experiment, in which we adjust four attention heads in middle layers in order to 'repair' the Colored Objects circuit and make it behave like the IOI circuit. In doing so, we boost accuracy from 49.6% to 93.7% on the Colored Objects task and explain most sources of error. The intervention affects downstream attention heads in specific ways predicted by their interactions in the IOI circuit, indicating that this subcircuit behavior is invariant to the different task inputs. Overall, our results provide evidence that it may yet be possible to explain large language models' behavior in terms of a relatively small number of interpretable task-general algorithmic building blocks and computational components. 1
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper49
- Assessing the Brittleness of Safety Alignment via Pruning and Low-Rank ModificationsBoyi Wei, Kaixuan Huang, Yangsibo Huang, Tinghao Xie 等ICML 2024 · 被引用 215 次
- What needs to go right for an induction head? A mechanistic study of in-context learning circuits and their formationAaditya K. Singh, Ted Moskovitz, Felix Hill, Stephanie C. Y. Chan 等ICML 2024 · 被引用 77 次
- LLM Circuit Analyses Are Consistent Across Training and ScaleCurt Tigges, Michael Hanna, Qinan Yu, Stella BidermanNeurIPS 2024 · 被引用 71 次
- Knowledge Circuits in Pretrained TransformersYunzhi Yao, Ningyu Zhang, Zekun Xi, Mengru Wang 等NeurIPS 2024 · 被引用 71 次
- Talking Heads: Understanding Inter-Layer Communication in Transformer Language ModelsJack Merullo, Carsten Eickhoff, Ellie PavlickNeurIPS 2024 · 被引用 49 次
它引用的顶会 Paper8
- Towards Automated Circuit Discovery for Mechanistic InterpretabilityArthur Conmy, Augustine N. Mavor-Parker, Aengus Lynch, Stefan Heimersheim 等NeurIPS 2023 · 被引用 861 次
- Investigating Gender Bias in Language Models Using Causal Mediation AnalysisJesse Vig, Sebastian Gehrmann, Yonatan Belinkov, Sharon Qian 等NeurIPS 2020 · 被引用 851 次
- Attention is Not Only a Weight: Analyzing Transformers with Vector NormsGoro Kobayashi, Tatsuki Kuribayashi, Sho Yokoi, Kentaro InuiEMNLP 2020 · 被引用 138 次
- Progress measures for grokking via mechanistic interpretabilityNeel Nanda, Lawrence Chan, Tom Lieberum, Jess Smith 等ICLR 2023 · 被引用 54 次
- Mass-Editing Memory in a TransformerKevin Meng, Arnab Sen Sharma, Alex J. Andonian, Yonatan Belinkov 等ICLR 2023 · 被引用 52 次
相关 Paper
- Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 SmallKevin Ro Wang, Alexandre Variengien, Arthur Conmy, Buck Shlegeris 等ICLR 2023 · 被引用 50 次
- Towards Interpretable Sequence Continuation: Analyzing Shared Circuits in Large Language ModelsMichael Lan, Philip Torr, Fazl BarezEMNLP 2024 · 被引用 1 次
- Beyond Components: Singular Vector-Based Interpretability of Transformer CircuitsAreeb Ahmad, Abhinav Joshi, Ashutosh ModiNeurIPS 2025 · 被引用 9 次
- Circuit Compositions: Exploring Modular Structures in Transformer-Based Language ModelsPhilipp Mondorf, Sondre Wold, Barbara PlankACL 2025 · 被引用 5 次
- Function Induction and Task Generalization: An Interpretability Study with Off-by-One AdditionQinyuan Ye, Robin Jia, Xiang RenICLR 2026 · 被引用 4 次
