Influence Paths for Characterizing Subject-Verb Number Agreement in LSTM Language Models
Kaiji Lu, Piotr Mardziel, Klas Leino, Matt Fredrikson, Anupam Datta
Abstract
LSTM-based recurrent neural networks are the state-of-the-art for many natural language processing (NLP) tasks. Despite their performance, it is unclear whether, or how, LSTMs learn structural features of natural languages such as subject-verb number agreement in English. Lacking this understanding, the generality of LSTMs on this task and their suitability for related tasks remains uncertain. Further, errors cannot be properly attributed to a lack of structural capability, training data omissions, or other exceptional faults. We introduce influence paths, a causal account of structural properties as carried by paths across gates and neurons of a recurrent neural network. The approach refines the notion of influence (the subject's grammatical number has influence on the grammatical number of the subsequent verb) into a set of gate-level or neuron-level paths. The set localizes and segments the concept (e.g., subject-verb agreement), its constituent elements (e.g., the subject), and related or interfering elements (e.g., attractors). We exemplify the methodology on a widely-studied multi-layer LSTM language model, demonstrating its accounting for subject-verb number agreement. The results offer both a finer and a more complete view of an LSTM's handling of this structural aspect of the English language than prior results based on diagnostic classifiers and ablation. Candidate Cell c1 1 Cell c 1 1 Candidate Cell c0 1 Hidden h 0 1 c 1 2 c 1 3 c 1 4 h 1 4 c 0 4 c1 4 c0 4 h 0 4 agreement s4(run) -s4(runs) grammatical number boys -boy+boys
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 5d161fdb-cafb-4344-925e-a9ffa590887eCited by top-tier papers4
- Influence Patterns for Explaining Information Flow in BERTKaiji Lu, Zifan Wang, Piotr Mardziel, Anupam DattaNeurIPS 2021 · 22 citations
- Generalizing Backpropagation for Gradient-Based InterpretabilityKevin Du, Lucas Torroba Hennigen, Niklas Stoehr, Alex Warstadt et al.ACL 2023 · 3 citations
- Discovering Influential Neuron Path in Vision TransformersYifan Wang, Yifei Liu, Yingdong Shi, Changming Li et al.ICLR 2025
- Causal Analysis of Syntactic Agreement Mechanisms in Neural Language ModelsMatthew Finlayson, Aaron Mueller, Sebastian Gehrmann, Stuart M. Shieber et al.ACL 2021
Builds on1
Related papers
- Different types of syntactic agreement recruit the same units within large language modelsDaria Kryvosheieva, Andrea Gregor de Varda, Evelina Fedorenko, Greta TuckuteACL 2026 · 3 citations
- Word Frequency Does Not Predict Grammatical Knowledge in Language ModelsCharles Yu, Ryan Sie, Nico Tedeschi, Leon BergenEMNLP 2020 · 6 citations
- An Analysis of the Utility of Explicit Negative Examples to Improve the Syntactic Abilities of Neural Language ModelsHiroshi Noji, Hiroya TakamuraACL 2020 · 12 citations
- A Comprehensive Comparison of Neural Networks as Cognitive Models of InflectionAdam Wiemerslage, Shiran Dudy, Katharina KannEMNLP 2022 · 3 citations
- Understanding or Memorizing? A Case Study of German Definite Articles in Language ModelsJonathan Drechsel, Erisa Bytyqi, Steffen HerboldACL 2026 · 1 citation
