Lower Perplexity is Not Always Human-Like
Tatsuki Kuribayashi, Yohei Oseki, Takumi Ito, Ryo Yoshida, Masayuki Asahara, Kentaro Inui
Abstract
In computational psycholinguistics, various language models have been evaluated against human reading behavior (e.g., eye movement) to build human-like computational models. However, most previous efforts have focused almost exclusively on English, despite the recent trend towards linguistic universal within the general community. In order to fill the gap, this paper investigates whether the established results in computational psycholinguistics can be generalized across languages. Specifically, we re-examine an established generalization -the lower perplexity a language model has, the more human-like the language model isin Japanese with typologically different structures from English. Our experiments demonstrate that this established generalization exhibits a surprising lack of universality; namely, lower perplexity is not always human-like. Moreover, this discrepancy between English and Japanese is further explored from the perspective of (non-)uniform information density. Overall, our results suggest that a crosslingual evaluation will be necessary to construct human-like computational models.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b2471398-a385-43f3-aee1-5ce433b88ed8Cited by top-tier papers13
- Context Limitations Make Neural Language Models More Human-LikeTatsuki Kuribayashi, Yohei Oseki, Ana Brassard, Kentaro InuiEMNLP 2022 · 28 citations
- Do We Truly Need So Many Samples? Multi-LLM Repeated Sampling Efficiently Scales Test-Time ComputeJianhao Chen, Zishuo Xun, Bocheng Zhou, Han Qi et al.AAAI 2026 · 18 citations
- Structural Priming Demonstrates Abstract Grammatical Representations in Multilingual Language ModelsJames A. Michaelov, Catherine Arnett, Tyler A. Chang, Ben BergenEMNLP 2023 · 6 citations
- Revisiting the Uniform Information Density HypothesisClara Meister, Tiago Pimentel, Patrick Haller, Lena A. Jäger et al.EMNLP 2021 · 4 citations
- A Spatio-Temporal Point Process for Fine-Grained Modeling of Reading BehaviorFrancesco Ignazio Re, Andreas Opedal, Glib Manaiev, Mario Giulianelli et al.ACL 2025 · 2 citations
Builds on4
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Extracting Training Data from Large Language ModelsNicholas Carlini, Florian Tramèr, Eric Wallace, Matthew Jagielski et al.USENIX Security 2021 · 2,866 citations
- If beam search is the answer, what was the question?Clara Meister, Ryan Cotterell, Tim VieiraEMNLP 2020 · 26 citations
- A Cognitive Regularizer for Language ModelingJason Wei, Clara Meister, Ryan CotterellACL 2021
Related papers
- Probing for Reading TimesEleftheria Tsipidi, Samuel Kiegeland, Francesco Ignazio Re, Tianyang Xu et al.ACL 2026
- Emergent Word Order Universals from Cognitively-Motivated Language ModelsTatsuki Kuribayashi, Ryo Ueda, Ryo Yoshida, Yohei Oseki et al.ACL 2024 · 1 citation
- Arrows of Time for Large Language ModelsVassilis Papadopoulos, Jérémie Wenger, Clément HonglerICML 2024 · 16 citations
- Measuring the Impact of (Psycho-)Linguistic and Readability Features and Their Spill Over Effects on the Prediction of Eye Movement PatternsDaniel Wiechmann, Elma KerzACL 2022 · 17 citations
- Anything Goes? A Crosslinguistic Study of (Im)possible Language Learning in LMsXiulin Yang, Tatsuya Aoyama, Yuekun Yao, Ethan WilcoxACL 2025 · 9 citations
