CausalGym: Benchmarking causal interpretability methods on linguistic tasks
Aryaman Arora, Dan Jurafsky, Christopher Potts
Abstract
Language models (LMs) have proven to be powerful tools for psycholinguistic research, but most prior work has focused on purely behavioural measures (e.g., surprisal comparisons). At the same time, research in model interpretability has begun to illuminate the abstract causal mechanisms shaping LM behavior. To help bring these strands of research closer together, we introduce CausalGym. We adapt and expand the Syntax-Gym suite of tasks to benchmark the ability of interpretability methods to causally affect model behaviour. To illustrate how CausalGym can be used, we study the pythia models (14M-6.9B) and assess the causal efficacy of a wide range of interpretability methods, including linear probing and distributed alignment search (DAS). We find that DAS outperforms the other methods, and so we use it to study the learning trajectory of two difficult linguistic phenomena in pythia-1b: negative polarity item licensing and filler-gap dependencies. Our analysis shows that the mechanism implementing both of these tasks is learned in discrete stages, not gradually.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers10
- ReFT: Representation Finetuning for Language ModelsZhengxuan Wu, Aryaman Arora, Zheng Wang, Atticus Geiger et al.NeurIPS 2024 · 233 citations
- The Non-Linear Representation Dilemma: Is Causal Abstraction Enough for Mechanistic Interpretability?Denis Sutter, Julian Minder, Thomas Hofmann, Tiago PimentelNeurIPS 2025 · 30 citations
- Mechanisms vs. Outcomes: Probing for Syntax Fails to Explain Performance on Targeted Syntactic EvaluationsAnanth Agarwal, Jasper Jian, Christopher D. Manning, Shikhar MurtyEMNLP 2025 · 5 citations
- Using Shapley interactions to understand how models use structureDivyansh Singhvi, Diganta Misra, Andrej Erkelens, Raghav Jain et al.ACL 2025 · 1 citation
- Causal Interventions Reveal Shared Structure Across English Filler-Gap ConstructionsSasha Boguraev, Christopher Potts, Kyle MahowaldEMNLP 2025
Builds on5
- Pythia: A Suite for Analyzing Large Language Models Across Training and ScalingStella Biderman, Hailey Schoelkopf, Quentin Gregory Anthony, Herbie Bradley et al.ICML 2023 · 1,822 citations
- Causal Abstractions of Neural NetworksAtticus Geiger, Hanson Lu, Thomas Icard, Christopher PottsNeurIPS 2021 · 516 citations
- LEACE: Perfect linear concept erasure in closed formNora Belrose, David Schneider-Joseph, Shauli Ravfogel, Ryan Cotterell et al.NeurIPS 2023 · 305 citations
- Sudden Drops in the Loss: Syntax Acquisition, Phase Transitions, and Simplicity Bias in MLMsAngelica Chen, Ravid Shwartz-Ziv, Kyunghyun Cho, Matthew L. Leavitt et al.ICLR 2024 · 119 citations
- Inducing Causal Structure for Interpretable Neural NetworksAtticus Geiger, Zhengxuan Wu, Hanson Lu, Josh Rozner et al.ICML 2022 · 104 citations
Related papers
- Fine-Grained Analysis of Shared Syntactic Mechanisms in Language ModelsRyoma Kumon, Hitomi YanakaACL 2026
- Interpretability at Scale: Identifying Causal Mechanisms in AlpacaZhengxuan Wu, Atticus Geiger, Thomas Icard, Christopher Potts et al.NeurIPS 2023 · 146 citations
- Internal Causal Mechanisms Robustly Predict Language Model Out-of-Distribution BehaviorsJing Huang, Junyi Tao, Thomas Icard, Diyi Yang et al.ICML 2025
- Bridging Internal Consistency and External Alignment: A Causal and Dynamic Interpretability Framework for LLM GenerationShuyao Xiao, Shengling Wang, Ke ChaoACL 2026
- MIB: A Mechanistic Interpretability BenchmarkAaron Mueller, Atticus Geiger, Sarah Wiegreffe, Dana Arad et al.ICML 2025
