Global Information Thresholding for Sufficient and Necessary Circuits
Jegyeong Cho
Abstract
We study the problem of extracting causal circuits-small edge-level subgraphs inside a trained network that are sufficient on their own and necessary to the model’s behavior under explicit error control. Prior work largely optimizes observational rankings or applies ad-hoc sparsification, which can sever paths, ignore inhibitory edges, and admit ``ghost" components that fail under intervention. We recast circuit discovery as information-constrained selection rather than ranking: a single global threshold chooses edges by their marginal contribution, combined with a null hypothesis-based statistical threshold to control family-wise errors. Edge scores are computed by rank-consistent attribution aligned to the task metric, stabilized with Fisher-diagonal variance normalization, projected to an edge coordinate system that preserves paths, and enforced with hard gates for interventional semantics. We propose an evaluation protocol that prioritizes sufficiency/necessity (CPR, CMD), editability, error rates, and standard ranking metrics. The result is a small, path-faithful circuit with reproducible selection criteria. Our motivation is to replace visually appealing heatmaps with interventional guarantees and explicit error control.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on10
- Towards Automated Circuit Discovery for Mechanistic InterpretabilityArthur Conmy, Augustine N. Mavor-Parker, Aengus Lynch, Stefan Heimersheim et al.NeurIPS 2023 · 861 citations
- Causal Abstractions of Neural NetworksAtticus Geiger, Hanson Lu, Thomas Icard, Christopher PottsNeurIPS 2021 · 516 citations
- Understanding Deep Networks via Extremal Perturbations and Smooth MasksRuth Fong, Mandela Patrick, Andrea VedaldiICCV 2019 · 480 citations
- Interpreting Graph Neural Networks for NLP With Differentiable Edge MaskingMichael Sejr Schlichtkrull, Nicola De Cao, Ivan TitovICLR 2021 · 287 citations
- Sparse Autoencoders Learn Monosemantic Features in Vision-Language ModelsMateusz Pach, Shyamgopal Karthik, Quentin Bouniot, Serge J. Belongie et al.NeurIPS 2025 · 79 citations
Related papers
- Mechanistic Interpretability as Statistical Estimation: A Variance AnalysisMaxime Méloux, François Portet, Maxime PeyrardICML 2026 · 13 citations
- Neural Response Interpretation Through the Lens of Critical PathwaysAshkan Khakzar, Soroosh Baselizadeh, Saurabh Khanduja, Christian Rupprecht et al.CVPR 2021
- Certified Circuits: Stability Guarantees for Mechanistic CircuitsAlaa Anani, Tobias Lorenz, Bernt Schiele, Mario Fritz et al.ICML 2026 · 3 citations
- IBCircuit: Towards Holistic Circuit Discovery with Information BottleneckTian Bian, Yifan Niu, Chaohao Yuan, Chengzhi Piao et al.ICML 2025
- Sparse Feature Circuits: Discovering and Editing Interpretable Causal Graphs in Language ModelsSamuel Marks, Can Rager, Eric J. Michaud, Yonatan Belinkov et al.ICLR 2025
