Towards Interpretable Sequence Continuation: Analyzing Shared Circuits in Large Language Models
Michael Lan, Philip Torr, Fazl Barez
摘要
While transformer models exhibit strong capabilities on linguistic tasks, their complex architectures make them difficult to interpret.Recent work has aimed to reverse engineer transformer models into human-readable representations called circuits that implement algorithmic functions.We extend this research by analyzing and comparing circuits for similar sequence continuation tasks, which include increasing sequences of Arabic numerals, number words, and months.By applying circuit interpretability analysis, we identify a key sub-circuit in both GPT-2 Small and Llama-2-7B responsible for detecting sequence members and for predicting the next member in a sequence.Our analysis reveals that semantically related sequences rely on shared circuit subgraphs with analogous roles.Additionally, we show that this sub-circuit has effects on various math-related prompts, such as on intervaled circuits, Spanish number word and months continuation, and natural language word problems.This mechanistic understanding of transformers is a critical step towards building more robust, aligned, and interpretable language models. 1 To encourage reuse and further development our code and datasets can be found here:
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- Optimal ablation for interpretabilityMaximilian Li, Lucas JansonNeurIPS 2024 · 被引用 32 次
- Query Circuits: Explaining How Language Models Answer User PromptsTung-Yu Wu, Fazl BarezICML 2026 · 被引用 1 次
- Shared Lexical Task Representations Explain Behavioral Variability In LLMsZhuonan Yang, Jacob Xiaochen Li, Francisco Velez, Eric Todd 等ICML 2026
- Mind's Eye: A Benchmark of Visual Abstraction, Transformation and Composition for Multimodal LLMsRohit Sinha, Aditya Sanjiv Kanade, Sai Srinivas Kancheti, Vineeth N. Balasubramanian 等ACL 2026
- Inside-Out: Measuring Generalization in Vision Transformers Through Inner WorkingsYunxiang Peng, Mengmeng Ma, Ziyu Yao, Xi PengCVPR 2026
它引用的顶会 Paper14
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Locating and Editing Factual Associations in GPTKevin Meng, David Bau, Alex Andonian, Yonatan BelinkovNeurIPS 2022 · 被引用 3,415 次
- Towards Automated Circuit Discovery for Mechanistic InterpretabilityArthur Conmy, Augustine N. Mavor-Parker, Aengus Lynch, Stefan Heimersheim 等NeurIPS 2023 · 被引用 861 次
- Investigating Gender Bias in Language Models Using Causal Mediation AnalysisJesse Vig, Sebastian Gehrmann, Yonatan Belinkov, Sharon Qian 等NeurIPS 2020 · 被引用 851 次
- Language Models Represent Space and TimeWes Gurnee, Max TegmarkICLR 2024 · 被引用 303 次
相关 Paper
- Circuit Compositions: Exploring Modular Structures in Transformer-Based Language ModelsPhilipp Mondorf, Sondre Wold, Barbara PlankACL 2025 · 被引用 5 次
- Transcoders find interpretable LLM feature circuitsJacob Dunefsky, Philippe Chlenski, Neel NandaNeurIPS 2024 · 被引用 222 次
- Sparse Feature Circuits: Discovering and Editing Interpretable Causal Graphs in Language ModelsSamuel Marks, Can Rager, Eric J. Michaud, Yonatan Belinkov 等ICLR 2025
- Circuit Component Reuse Across Tasks in Transformer Language ModelsJack Merullo, Carsten Eickhoff, Ellie PavlickICLR 2024 · 被引用 108 次
- Towards Universality: Studying Mechanistic Similarity Across Language Model ArchitecturesJunxuan Wang, Xuyang Ge, Wentao Shu, Qiong Tang 等ICLR 2025
