Information Locality as an Inductive Bias for Neural Language Models
Taiga Someya, Anej Svete, Brian DuSell, Timothy J. O'Donnell, Mario Giulianelli, Ryan Cotterell
摘要
Inductive biases are inherent in every machine learning system, shaping how models generalize from finite data. In the case of neural language models (LMs), debates persist as to whether these biases align with or diverge from human processing constraints. To address this issue, we propose a quantitative framework that allows for controlled investigations into the nature of these biases. Within our framework, we introduce m-local entropy-an informationtheoretic measure derived from average lossycontext surprisal-that captures the local uncertainty of a language by quantifying how effectively the m -1 preceding symbols disambiguate the next symbol. In experiments on both perturbed natural language corpora and languages defined by probabilistic finite-state automata (PFSAs), we show that languages with higher m-local entropy are more difficult for Transformer and LSTM LMs to learn. These results suggest that neural LMs, much like humans, are highly sensitive to the local statistical structure of a language. https://github.com/rycolab/ lm-inductive-bias * This research was conducted while visiting ETH Zürich.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Causally Evaluating the Learnability of Formal Language TasksVésteinn Snæbjarnarson, Anej Svete, Josef Valvoda, Reda Boumasmoud 等ICML 2026
- Vocabulary Shapes Cross-Lingual Variation of Word-Order Learnability in Language ModelsJonas Mayer Martins, Jaap Jumelet, Viola Priesemann, Lisa BeinbornACL 2026
它引用的顶会 Paper5
- A Systematic Assessment of Syntactic Generalization in Neural Language ModelsJennifer Hu, Jon Gauthier, Peng Qian, Ethan Wilcox 等ACL 2020 · 被引用 124 次
- How poor is the stimulus? Evaluating hierarchical generalization in neural networks trained on child-directed speechAditya Yedetore, Tal Linzen, Robert Frank, R. Thomas McCoyACL 2023 · 被引用 17 次
- Mission: Impossible Language ModelsJulie Kallini, Isabel Papadimitriou, Richard Futrell, Kyle Mahowald 等ACL 2024 · 被引用 15 次
- Training Neural Networks as Recognizers of Formal LanguagesAlexandra Butoi, Ghazal Khalighinejad, Anej Svete, Josef Valvoda 等ICLR 2025
- What Languages are Easy to Language-Model? A Perspective from Learning Probabilistic Regular LanguagesNadav Borenstein, Anej Svete, Robin Chan, Josef Valvoda 等ACL 2024
相关 Paper
- Examining the Inductive Bias of Neural Language Models with Artificial LanguagesJennifer C. White, Ryan CotterellACL 2021
- Anything Goes? A Crosslinguistic Study of (Im)possible Language Learning in LMsXiulin Yang, Tatsuya Aoyama, Yuekun Yao, Ethan WilcoxACL 2025 · 被引用 9 次
- Context Limitations Make Neural Language Models More Human-LikeTatsuki Kuribayashi, Yohei Oseki, Ana Brassard, Kentaro InuiEMNLP 2022 · 被引用 28 次
- Arrows of Time for Large Language ModelsVassilis Papadopoulos, Jérémie Wenger, Clément HonglerICML 2024 · 被引用 16 次
- What they do when in doubt: a study of inductive biases in seq2seq learnersEugene Kharitonov, Rahma ChaabouniICLR 2021 · 被引用 29 次
