Causal Interventions Reveal Shared Structure Across English Filler-Gap Constructions
Sasha Boguraev, Christopher Potts, Kyle Mahowald
摘要
Language Models (LMs) have emerged as powerful sources of evidence for linguists seeking to develop theories of syntax. In this paper, we argue that causal interpretability methods, applied to LMs, can greatly enhance the value of such evidence by helping us characterize the abstract mechanisms that LMs learn to use. Our empirical focus is a set of English filler-gap dependency constructions (e.g., questions, relative clauses). Linguistic theories largely agree that these constructions share many properties. Using experiments based in Distributed Interchange Interventions, we show that LMs converge on similar abstract analyses of these constructions. These analyses also reveal previously overlooked factorsrelating to frequency, filler type, and surrounding context -that could motivate changes to standard linguistic theory. Overall, these results suggest that mechanistic, internal analyses of LMs can push linguistic theory forward. https://github.com/SashaBoguraev/ causal-filler-gap
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Different types of syntactic agreement recruit the same units within large language modelsDaria Kryvosheieva, Andrea Gregor de Varda, Evelina Fedorenko, Greta TuckuteACL 2026 · 被引用 3 次
- Fine-Grained Analysis of Shared Syntactic Mechanisms in Language ModelsRyoma Kumon, Hitomi YanakaACL 2026
它引用的顶会 Paper9
- Locating and Editing Factual Associations in GPTKevin Meng, David Bau, Alex Andonian, Yonatan BelinkovNeurIPS 2022 · 被引用 3,415 次
- Pythia: A Suite for Analyzing Large Language Models Across Training and ScalingStella Biderman, Hailey Schoelkopf, Quentin Gregory Anthony, Herbie Bradley 等ICML 2023 · 被引用 1,822 次
- Causal Abstractions of Neural NetworksAtticus Geiger, Hanson Lu, Thomas Icard, Christopher PottsNeurIPS 2021 · 被引用 516 次
- Interpretability at Scale: Identifying Causal Mechanisms in AlpacaZhengxuan Wu, Atticus Geiger, Thomas Icard, Christopher Potts 等NeurIPS 2023 · 被引用 146 次
- A Systematic Assessment of Syntactic Generalization in Neural Language ModelsJennifer Hu, Jon Gauthier, Peng Qian, Ethan Wilcox 等ACL 2020 · 被引用 124 次
相关 Paper
- CausalGym: Benchmarking causal interpretability methods on linguistic tasksAryaman Arora, Dan Jurafsky, Christopher PottsACL 2024 · 被引用 3 次
- Mind the Gap: How BabyLMs Learn Filler-Gap DependenciesChi-Yun Chang, Xueyang Huang, Humaira Nasir, Shane Storks 等EMNLP 2025
- LLM Interpretability with Identifiable Temporal-Instantaneous RepresentationXiangchen Song, Jiaqi Sun, Zijian Li, Yujia Zheng 等NeurIPS 2025 · 被引用 6 次
- Bridging Internal Consistency and External Alignment: A Causal and Dynamic Interpretability Framework for LLM GenerationShuyao Xiao, Shengling Wang, Ke ChaoACL 2026
- Internal Causal Mechanisms Robustly Predict Language Model Out-of-Distribution BehaviorsJing Huang, Junyi Tao, Thomas Icard, Diyi Yang 等ICML 2025
