Lune

NeurIPS2021Top-tier venue

Causal Abstractions of Neural Networks

Atticus Geiger, Hanson Lu, Thomas Icard, Christopher Potts

2021Year
516Citations
109Top-tier citations

Abstract

Structural analysis methods (e.g., probing and feature attribution) are increasingly important tools for neural network analysis. We propose a new structural analysis method grounded in a formal theory of causal abstraction that provides rich characterizations of model-internal representations and their roles in input/output behavior. In this method, neural representations are aligned with variables in interpretable causal models, and then interchange interventions are used to experimentally verify that the neural representations have the causal properties of their aligned variables. We apply this method in a case study to analyze neural models trained on Multiply Quantified Natural Language Inference (MQNLI) corpus, a highly complex NLI dataset that was constructed with a tree-structured natural logic causal model. We discover that a BERT-based model with state-of-the-art performance successfully realizes parts of the natural logic model's causal structure, whereas a simpler baseline model fails to show any such structure, demonstrating that BERT representations encode the compositional structure of MQNLI. * equal contribution 2 We provide tools for causal abstraction analysis at

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

lune papers fulltext aa055249-16fd-4091-88ff-8e23c7440d2b

Cited by top-tier papers109

Ask how each one uses it

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines