Circuit Compositions: Exploring Modular Structures in Transformer-Based Language Models
Philipp Mondorf, Sondre Wold, Barbara Plank
Abstract
A fundamental question in interpretability research is to what extent neural networks, particularly language models, implement reusable functions through subnetworks that can be composed to perform more complex tasks. Recent advances in mechanistic interpretability have made progress in identifying , which represent the minimal computational subgraphs responsible for a model's behavior on specific tasks. However, most studies focus on identifying circuits for individual tasks without investigating how functionally similar circuits to each other. To address this gap, we study the modularity of neural networks by analyzing circuits for highly compositional subtasks within a transformer-based language model. Specifically, given a probabilistic context-free grammar, we identify and compare circuits responsible for ten modular string-edit operations. Our results indicate that functionally similar circuits exhibit both notable node overlap and cross-task faithfulness. Moreover, we demonstrate that the circuits identified can be reused and combined through set operations to represent more complex functional model capabilities.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 30a95f4d-6887-408d-9d80-366c4bec06f3Cited by top-tier papers11
- Same Task, Different Circuits: Disentangling Modality-Specific Mechanisms in VLMsYaniv Nikankin, Dana Arad, Yossi Gandelsman, Yonatan BelinkovNeurIPS 2025 · 37 citations
- Sheaf Discovery with Joint Computation Graph Pruning and Flexible GranularityLei Yu, Jingcheng Niu, Zining Zhu, Xi Chen et al.EMNLP 2025 · 11 citations
- The Validation Gap: A Mechanistic Analysis of How Language Models Compute Arithmetic but Fail to Validate ItLeonardo Bertolazzi, Philipp Mondorf, Barbara Plank, Raffaella BernardiEMNLP 2025 · 8 citations
- On Relation-Specific Neurons in Large Language ModelsYihong Liu, Runsheng Chen, Lea Hirlimann, Ahmad Dawar Hakimi et al.EMNLP 2025
- Interpreting and Enhancing Emotional Circuits in Large Vision-Language Models via Cross-Modal Information FlowChengsheng Zhang, Chenghao Sun, Zhining Xie, Xinmei TianICML 2026
Builds on16
- Locating and Editing Factual Associations in GPTKevin Meng, David Bau, Alex Andonian, Yonatan BelinkovNeurIPS 2022 · 3,415 citations
- Towards Automated Circuit Discovery for Mechanistic InterpretabilityArthur Conmy, Augustine N. Mavor-Parker, Aengus Lynch, Stefan Heimersheim et al.NeurIPS 2023 · 861 citations
- Investigating Gender Bias in Language Models Using Causal Mediation AnalysisJesse Vig, Sebastian Gehrmann, Yonatan Belinkov, Sharon Qian et al.NeurIPS 2020 · 851 citations
- Thinking Like TransformersGail Weiss, Yoav Goldberg, Eran YahavICML 2021 · 183 citations
- Winning the Lottery with Continuous SparsificationPedro Savarese, Hugo Silva, Michael MaireNeurIPS 2020 · 162 citations
Related papers
- Towards Interpretable Sequence Continuation: Analyzing Shared Circuits in Large Language ModelsMichael Lan, Philip Torr, Fazl BarezEMNLP 2024 · 1 citation
- Sparse Feature Circuits: Discovering and Editing Interpretable Causal Graphs in Language ModelsSamuel Marks, Can Rager, Eric J. Michaud, Yonatan Belinkov et al.ICLR 2025
- Towards Global-level Mechanistic Interpretability: A Perspective of Modular Circuits of Large Language ModelsYinhan He, Wendy Zheng, Yushun Dong, Yaochen Zhu et al.ICML 2025
- Efficient Automated Circuit Discovery in Transformers using Contextual DecompositionAliyah R. Hsu, Georgia Zhou, Yeshwanth Cherapanamjeri, Yaxuan Huang et al.ICLR 2025
- Beyond Components: Singular Vector-Based Interpretability of Transformer CircuitsAreeb Ahmad, Abhinav Joshi, Ashutosh ModiNeurIPS 2025 · 9 citations
