Sheaf Discovery with Joint Computation Graph Pruning and Flexible Granularity
Lei Yu, Jingcheng Niu, Zining Zhu, Xi Chen, Gerald Penn
摘要
In this paper, we introduce DiscoGP, a novel framework for extracting self-contained modular units, or sheaves, within neural language models (LMs). Sheaves extend the concept of functional circuits, a unit widely explored in interpretability research, by considering not only subsets of edges in an LM's computation graph but also the model's weight parameters. Our framework identifies sheaves through a gradient-based pruning algorithm that operates on both of these in such a way that reduces the original LM to a sparse skeleton that preserves certain core capabilities. Experimental results demonstrate that, across a range of linguistic and reasoning tasks, DiscoGP extracts sheaves that preserve 93-100% of a model's performance on the identified task while comprising only 1-7% of the original weights and connections. Furthermore, our analysis reveals that, compared to previously identified LM circuits, the sheaves discovered by DiscoGP exhibit superior modularity and functional fidelity. Extending our method to the neuron level also unveils novel insights into the inner workings of LLMs. 1 * Equal contribution. 1 The code and results of DiscoGP are available online: https://github.com/frankniujc/disco_gp .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Mechanistic Interpretability as Statistical Estimation: A Variance AnalysisMaxime Méloux, François Portet, Maxime PeyrardICML 2026 · 被引用 13 次
- All Circuits Lead to Rome: Rethinking Functional Anisotropy in Circuit and Sheaf Discovery for LLMsXi Chen, Mingyu Jin, Jingcheng (Frank) Niu, Yutong Yin 等ICML 2026 · 被引用 5 次
- Granular Concept Circuits: Toward a Fine-Grained Circuit Discovery for Concept RepresentationsDahee Kwon, Sehyun Lee, Jaesik ChoiICCV 2025
它引用的顶会 Paper30
- Locating and Editing Factual Associations in GPTKevin Meng, David Bau, Alex Andonian, Yonatan BelinkovNeurIPS 2022 · 被引用 3,415 次
- Sparse Autoencoders Find Highly Interpretable Features in Language ModelsRobert Huben, Hoagy Cunningham, Logan Riggs Smith, Aidan Ewart 等ICLR 2024 · 被引用 1,072 次
- Towards Automated Circuit Discovery for Mechanistic InterpretabilityArthur Conmy, Augustine N. Mavor-Parker, Aengus Lynch, Stefan Heimersheim 等NeurIPS 2023 · 被引用 861 次
- Investigating Gender Bias in Language Models Using Causal Mediation AnalysisJesse Vig, Sebastian Gehrmann, Yonatan Belinkov, Sharon Qian 等NeurIPS 2020 · 被引用 851 次
- Causal Abstractions of Neural NetworksAtticus Geiger, Hanson Lu, Thomas Icard, Christopher PottsNeurIPS 2021 · 被引用 516 次
相关 Paper
- Weight-sparse transformers have interpretable circuitsLeo Gao, Achyuta Rajaram, Jacob Coxon, Soham Govande 等ICML 2026
- Discovering Decoupled Functional Modules in Large Language ModelsYanke Yu, Jin Li, Ying Sun, Ping Li 等AAAI 2026
- Extracting Interpretable Task-Specific Circuits from Large Language Models for Faster InferenceJorge García-Carrasco, Alejandro Maté, Juan TrujilloAAAI 2025 · 被引用 3 次
- Sparse Feature Circuits: Discovering and Editing Interpretable Causal Graphs in Language ModelsSamuel Marks, Can Rager, Eric J. Michaud, Yonatan Belinkov 等ICLR 2025
- Language Model Circuits Are Sparse in the Neuron BasisAryaman Arora, Zhengxuan Wu, Jacob Steinhardt, Sarah SchwettmannICML 2026
