Influence Paths for Characterizing Subject-Verb Number Agreement in LSTM Language Models
Kaiji Lu, Piotr Mardziel, Klas Leino, Matt Fredrikson, Anupam Datta
摘要
LSTM-based recurrent neural networks are the state-of-the-art for many natural language processing (NLP) tasks. Despite their performance, it is unclear whether, or how, LSTMs learn structural features of natural languages such as subject-verb number agreement in English. Lacking this understanding, the generality of LSTMs on this task and their suitability for related tasks remains uncertain. Further, errors cannot be properly attributed to a lack of structural capability, training data omissions, or other exceptional faults. We introduce influence paths, a causal account of structural properties as carried by paths across gates and neurons of a recurrent neural network. The approach refines the notion of influence (the subject's grammatical number has influence on the grammatical number of the subsequent verb) into a set of gate-level or neuron-level paths. The set localizes and segments the concept (e.g., subject-verb agreement), its constituent elements (e.g., the subject), and related or interfering elements (e.g., attractors). We exemplify the methodology on a widely-studied multi-layer LSTM language model, demonstrating its accounting for subject-verb number agreement. The results offer both a finer and a more complete view of an LSTM's handling of this structural aspect of the English language than prior results based on diagnostic classifiers and ablation. Candidate Cell c1 1 Cell c 1 1 Candidate Cell c0 1 Hidden h 0 1 c 1 2 c 1 3 c 1 4 h 1 4 c 0 4 c1 4 c0 4 h 0 4 agreement s4(run) -s4(runs) grammatical number boys -boy+boys
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Influence Patterns for Explaining Information Flow in BERTKaiji Lu, Zifan Wang, Piotr Mardziel, Anupam DattaNeurIPS 2021 · 被引用 22 次
- Generalizing Backpropagation for Gradient-Based InterpretabilityKevin Du, Lucas Torroba Hennigen, Niklas Stoehr, Alex Warstadt 等ACL 2023 · 被引用 3 次
- Discovering Influential Neuron Path in Vision TransformersYifan Wang, Yifei Liu, Yingdong Shi, Changming Li 等ICLR 2025
- Causal Analysis of Syntactic Agreement Mechanisms in Neural Language ModelsMatthew Finlayson, Aaron Mueller, Sebastian Gehrmann, Stuart M. Shieber 等ACL 2021
它引用的顶会 Paper1
相关 Paper
- Different types of syntactic agreement recruit the same units within large language modelsDaria Kryvosheieva, Andrea Gregor de Varda, Evelina Fedorenko, Greta TuckuteACL 2026 · 被引用 3 次
- Word Frequency Does Not Predict Grammatical Knowledge in Language ModelsCharles Yu, Ryan Sie, Nico Tedeschi, Leon BergenEMNLP 2020 · 被引用 6 次
- An Analysis of the Utility of Explicit Negative Examples to Improve the Syntactic Abilities of Neural Language ModelsHiroshi Noji, Hiroya TakamuraACL 2020 · 被引用 12 次
- A Comprehensive Comparison of Neural Networks as Cognitive Models of InflectionAdam Wiemerslage, Shiran Dudy, Katharina KannEMNLP 2022 · 被引用 3 次
- Understanding or Memorizing? A Case Study of German Definite Articles in Language ModelsJonathan Drechsel, Erisa Bytyqi, Steffen HerboldACL 2026 · 被引用 1 次
