Lune

ICML2022Top-tier venue

Neuron Dependency Graphs: A Causal Abstraction of Neural Networks

Yaojie Hu, Jin Tian

2022Year
8Citations

Abstract

We discover that neural networks exhibit approximate logical dependencies among neurons, and we introduce Neuron Dependency Graphs (NDG) that extract and present them as directed graphs. In an NDG, each node corresponds to the boolean activation value of a neuron, and each edge models an approximate logical implication from one node to another. We show that the logical dependencies extracted from the training dataset generalize well to the test set. In addition to providing symbolic explanations to the neural network's internal structure, NDGs can represent a Structural Causal Model. We empirically show that an NDG is a causal abstraction of the corresponding neural network that "unfolds" the same way under causal interventions using the theory by Geiger et al. (2021a). Code is available at https://github.com/phimachine/ndg .

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

lune papers fulltext 9cefe141-8195-49aa-aa9f-d18734cf5ddc

Builds on5

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines