Information Locality as an Inductive Bias for Neural Language Models
Taiga Someya, Anej Svete, Brian DuSell, Timothy J. O'Donnell, Mario Giulianelli, Ryan Cotterell
Abstract
Inductive biases are inherent in every machine learning system, shaping how models generalize from finite data. In the case of neural language models (LMs), debates persist as to whether these biases align with or diverge from human processing constraints. To address this issue, we propose a quantitative framework that allows for controlled investigations into the nature of these biases. Within our framework, we introduce m-local entropy-an informationtheoretic measure derived from average lossycontext surprisal-that captures the local uncertainty of a language by quantifying how effectively the m -1 preceding symbols disambiguate the next symbol. In experiments on both perturbed natural language corpora and languages defined by probabilistic finite-state automata (PFSAs), we show that languages with higher m-local entropy are more difficult for Transformer and LSTM LMs to learn. These results suggest that neural LMs, much like humans, are highly sensitive to the local statistical structure of a language. https://github.com/rycolab/ lm-inductive-bias * This research was conducted while visiting ETH Zürich.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 79e5c9a8-9c2e-443f-8b83-0b0be88a2ebeCited by top-tier papers2
- Causally Evaluating the Learnability of Formal Language TasksVésteinn Snæbjarnarson, Anej Svete, Josef Valvoda, Reda Boumasmoud et al.ICML 2026
- Vocabulary Shapes Cross-Lingual Variation of Word-Order Learnability in Language ModelsJonas Mayer Martins, Jaap Jumelet, Viola Priesemann, Lisa BeinbornACL 2026
Builds on5
- A Systematic Assessment of Syntactic Generalization in Neural Language ModelsJennifer Hu, Jon Gauthier, Peng Qian, Ethan Wilcox et al.ACL 2020 · 124 citations
- How poor is the stimulus? Evaluating hierarchical generalization in neural networks trained on child-directed speechAditya Yedetore, Tal Linzen, Robert Frank, R. Thomas McCoyACL 2023 · 17 citations
- Mission: Impossible Language ModelsJulie Kallini, Isabel Papadimitriou, Richard Futrell, Kyle Mahowald et al.ACL 2024 · 15 citations
- Training Neural Networks as Recognizers of Formal LanguagesAlexandra Butoi, Ghazal Khalighinejad, Anej Svete, Josef Valvoda et al.ICLR 2025
- What Languages are Easy to Language-Model? A Perspective from Learning Probabilistic Regular LanguagesNadav Borenstein, Anej Svete, Robin Chan, Josef Valvoda et al.ACL 2024
Related papers
- Examining the Inductive Bias of Neural Language Models with Artificial LanguagesJennifer C. White, Ryan CotterellACL 2021
- Anything Goes? A Crosslinguistic Study of (Im)possible Language Learning in LMsXiulin Yang, Tatsuya Aoyama, Yuekun Yao, Ethan WilcoxACL 2025 · 9 citations
- Context Limitations Make Neural Language Models More Human-LikeTatsuki Kuribayashi, Yohei Oseki, Ana Brassard, Kentaro InuiEMNLP 2022 · 28 citations
- Arrows of Time for Large Language ModelsVassilis Papadopoulos, Jérémie Wenger, Clément HonglerICML 2024 · 16 citations
- What they do when in doubt: a study of inductive biases in seq2seq learnersEugene Kharitonov, Rahma ChaabouniICLR 2021 · 29 citations
