Emergent Word Order Universals from Cognitively-Motivated Language Models
Tatsuki Kuribayashi, Ryo Ueda, Ryo Yoshida, Yohei Oseki, Ted Briscoe, Timothy Baldwin
Abstract
The world's languages exhibit certain so-called typological or implicational universals; for example, Subject-Object-Verb (SOV) languages typically use postpositions. Explaining the source of such biases is a key goal of linguistics. We study word-order universals through a computational simulation with language models (LMs). Our experiments show that typologically-typical word orders tend to have lower perplexity estimated by LMs with cognitively plausible biases: syntactic biases, specific parsing strategies, and memory limitations. This suggests that the interplay of cognitive biases and predictability (perplexity) can explain many aspects of word-order universals. It also showcases the advantage of cognitively-motivated LMs, typically employed in cognitive modeling, in the simulation of language universals. https://github.com/kuribayashi4/ word-order-universals-cogLM ory (Smith and Levy, 2013), but such a logarithmic conversion did not alter our findings ( §7). Thus, we use PPL, following White and Cotterell (2021) .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 105082f6-dbb1-4c11-aaa6-7bb3f7e579caCited by top-tier papers2
- On the Effect of Hyperparameters in Language Modeling for Computational LinguisticsRuoxi Ning, Yongpeng Zhu, Qingcheng Zeng, Tatsuki Kuribayashi et al.ACL 2026
- Which Word Orders Facilitate Length Generalization in LMs? An Investigation with GCG-Based Artificial LanguagesNadine El-Naggar, Tatsuki Kuribayashi, Ted BriscoeEMNLP 2025
Builds on10
- Neural Networks and the Chomsky HierarchyGrégoire Delétang, Anian Ruoss, Jordi Grau-Moya, Tim Genewein et al.ICLR 2023 · 45 citations
- Emergent Communication: Generalization and Overfitting in Lewis GamesMathieu Rita, Corentin Tallec, Paul Michel, Jean-Bastien Grill et al.NeurIPS 2022 · 41 citations
- What they do when in doubt: a study of inductive biases in seq2seq learnersEugene Kharitonov, Rahma ChaabouniICLR 2021 · 29 citations
- Context Limitations Make Neural Language Models More Human-LikeTatsuki Kuribayashi, Yohei Oseki, Ana Brassard, Kentaro InuiEMNLP 2022 · 28 citations
- Mission: Impossible Language ModelsJulie Kallini, Isabel Papadimitriou, Richard Futrell, Kyle Mahowald et al.ACL 2024 · 15 citations
Related papers
- Examining the Inductive Bias of Neural Language Models with Artificial LanguagesJennifer C. White, Ryan CotterellACL 2021
- Anything Goes? A Crosslinguistic Study of (Im)possible Language Learning in LMsXiulin Yang, Tatsuya Aoyama, Yuekun Yao, Ethan WilcoxACL 2025 · 9 citations
- Lower Perplexity is Not Always Human-LikeTatsuki Kuribayashi, Yohei Oseki, Takumi Ito, Ryo Yoshida et al.ACL 2021
- Language Models as an Alternative Evaluator of Word Order Hypotheses: A Case Study in JapaneseTatsuki Kuribayashi, Takumi Ito, Jun Suzuki, Kentaro InuiACL 2020 · 2 citations
- Vocabulary Shapes Cross-Lingual Variation of Word-Order Learnability in Language ModelsJonas Mayer Martins, Jaap Jumelet, Viola Priesemann, Lisa BeinbornACL 2026
