Fine-Grained Analysis of Shared Syntactic Mechanisms in Language Models
Ryoma Kumon, Hitomi Yanaka
摘要
While language models demonstrate sophisticated syntactic capabilities, the extent to which their internal mechanisms align with crossconstructional principles studied in linguistics remains poorly understood. This study investigates whether models employ shared neural mechanisms across different syntactic constructions by applying causal interpretability methods at a granular level. Focusing on filler-gap dependencies and negative polarity item (NPI) licensing, we utilize activation patching to identify the functional roles of specific attention heads and MLP blocks. Our results reveal a highly localized and shared mechanism for filler-gap dependencies located in the early to middle layers, whereas NPI processing exhibits no such unified mechanism. Furthermore, we find that these mechanisms identified by activation patching generalize to out-of-distribution, while distributed alignment search, a supervised interpretability method, is susceptible to overfitting on narrow linguistic distributions. Finally, we validate our findings by demonstrating that the manipulation of the identified components improves model performance on acceptability judgment benchmarks. 1
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper9
- Investigating Gender Bias in Language Models Using Causal Mediation AnalysisJesse Vig, Sebastian Gehrmann, Yonatan Belinkov, Sharon Qian 等NeurIPS 2020 · 被引用 851 次
- Causal Abstractions of Neural NetworksAtticus Geiger, Hanson Lu, Thomas Icard, Christopher PottsNeurIPS 2021 · 被引用 516 次
- A Systematic Assessment of Syntactic Generalization in Neural Language ModelsJennifer Hu, Jon Gauthier, Peng Qian, Ethan Wilcox 等ACL 2020 · 被引用 124 次
- Probing for the Usage of Grammatical NumberKarim Lasri, Tiago Pimentel, Alessandro Lenci, Thierry Poibeau 等ACL 2022 · 被引用 72 次
- CausalGym: Benchmarking causal interpretability methods on linguistic tasksAryaman Arora, Dan Jurafsky, Christopher PottsACL 2024 · 被引用 3 次
相关 Paper
- Causal Interventions Reveal Shared Structure Across English Filler-Gap ConstructionsSasha Boguraev, Christopher Potts, Kyle MahowaldEMNLP 2025
- How Language Models Process NegationZhejian Zhou, Tianyi Zhou, Robin Jia, Jonathan MayICML 2026
- Internal Causal Mechanisms Robustly Predict Language Model Out-of-Distribution BehaviorsJing Huang, Junyi Tao, Thomas Icard, Diyi Yang 等ICML 2025
- Bridging the Language Gap: Uncovering and Aligning Shared Circuits for Multi-Hop Reasoning in Multilingual LLMsChenghao Sun, Zhen Huang, Yonggang Zhang, Xinmei Tian 等AAAI 2026
- Interpreting Language Models with Contrastive ExplanationsKayo Yin, Graham NeubigEMNLP 2022 · 被引用 32 次
