Understanding or Memorizing? A Case Study of German Definite Articles in Language Models
Jonathan Drechsel, Erisa Bytyqi, Steffen Herbold
摘要
Language models perform well on grammatical agreement, but it is unclear whether this reflects rule-based generalization or memorization. We study this question for German definite singular articles, whose forms depend on gender and case. Using GRADIEND, a gradient-based interpretability method, we learn parameter update directions for gender-case specific article transitions. We find that updates learned for a specific gender-case article transition frequently affect unrelated gender-case settings, with substantial overlap among the most affected neurons across settings. These results argue against a strictly rule-based encoding of German definite articles, indicating that models at least partly rely on memorized associations rather than abstract grammatical rules.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper5
- Quantifying Memorization Across Neural Language ModelsNicholas Carlini, Daphne Ippolito, Matthew Jagielski, Katherine Lee 等ICLR 2023 · 被引用 158 次
- Counting the Bugs in ChatGPT's Wugs: A Multilingual Investigation into the Morphological Capabilities of a Large Language ModelLeonie Weissweiler, Valentin Hofmann, Anjali Kantharuban, Anna Cai 等EMNLP 2023 · 被引用 10 次
- LinguaLens: Towards Interpreting Linguistic Mechanisms of Large Language Models via Sparse Auto-EncoderYi Jing, Zijun Yao, Hongzhu Guo, Lingxu Ran 等EMNLP 2025 · 被引用 7 次
- GRADIEND: Feature Learning within Neural Networks Exemplified through BiasesJonathan Drechsel, Steffen HerboldICLR 2026 · 被引用 2 次
- Causal Analysis of Syntactic Agreement Mechanisms in Neural Language ModelsMatthew Finlayson, Aaron Mueller, Sebastian Gehrmann, Stuart M. Shieber 等ACL 2021
相关 Paper
- Influence Paths for Characterizing Subject-Verb Number Agreement in LSTM Language ModelsKaiji Lu, Piotr Mardziel, Klas Leino, Matt Fredrikson 等ACL 2020 · 被引用 8 次
- Generalizing Backpropagation for Gradient-Based InterpretabilityKevin Du, Lucas Torroba Hennigen, Niklas Stoehr, Alex Warstadt 等ACL 2023 · 被引用 3 次
- Frequency Effects on Syntactic Rule Learning in TransformersJason Wei, Dan Garrette, Tal Linzen, Ellie PavlickEMNLP 2021 · 被引用 40 次
- Different types of syntactic agreement recruit the same units within large language modelsDaria Kryvosheieva, Andrea Gregor de Varda, Evelina Fedorenko, Greta TuckuteACL 2026 · 被引用 3 次
- Investigating Gender Bias in Language Models Using Causal Mediation AnalysisJesse Vig, Sebastian Gehrmann, Yonatan Belinkov, Sharon Qian 等NeurIPS 2020 · 被引用 851 次
