Reverse-Engineering the Reader
Samuel Kiegeland, Ethan Wilcox, Afra Amini, David Robert Reich, Ryan Cotterell
Abstract
Numerous previous studies have sought to determine to what extent language models, pretrained on natural language text, can serve as useful models of human cognition. In this paper, we are interested in the opposite question: whether we can directly optimize a language model to be a useful cognitive model by aligning it to human psychometric data. To achieve this, we introduce a novel alignment technique in which we fine-tune a language model to implicitly optimize the parameters of a linear regressor that directly predicts humans' reading times of in-context linguistic units, e.g., phonemes, morphemes, or words, using surprisal estimates derived from the language model. Using words as a test case, we evaluate our technique across multiple model sizes and datasets and find that it improves language models' psychometric predictive power. However, we find an inverse relationship between psychometric power and a model's performance on downstream NLP tasks as well as its perplexity on held-out test data. While this latter trend has been observed before (Oh et al., 2022; Shain et al., 2024) , we are the first to induce it by manipulating a model's alignment to psychometric data.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 90f1aef1-366d-4509-9503-e1fbfc8f66d9Cited by top-tier papers1
Ask how each one uses itBuilds on8
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning et al.NeurIPS 2023 · 10,924 citations
- Pythia: A Suite for Analyzing Large Language Models Across Training and ScalingStella Biderman, Hailey Schoelkopf, Quentin Gregory Anthony, Herbie Bradley et al.ICML 2023 · 1,822 citations
- Learning to summarize with human feedbackNisan Stiennon, Long Ouyang, Jeffrey Wu, Daniel M. Ziegler et al.NeurIPS 2020 · 124 citations
- Revisiting the Uniform Information Density HypothesisClara Meister, Tiago Pimentel, Patrick Haller, Lena A. Jäger et al.EMNLP 2021 · 4 citations
Related papers
- Probing for Reading TimesEleftheria Tsipidi, Samuel Kiegeland, Francesco Ignazio Re, Tianyang Xu et al.ACL 2026
- Temperature-scaling surprisal estimates improve fit to human reading times - but does it do so for the "right reasons"?Tong Liu, Iza Skrjanec, Vera DembergACL 2024 · 3 citations
- Language models and brains align due to more than next-word prediction and word-level informationGabriele Merlin, Mariya TonevaEMNLP 2024 · 2 citations
- From Human Reading to NLM Understanding: Evaluating the Role of Eye-Tracking Data in Encoder-Based ModelsLuca Dini, Lucia Domenichelli, Dominique Brunato, Felice Dell'OrlettaACL 2025
- Dual Alignment Between Language Model Layers and Human Sentence ProcessingTatsuki Kuribayashi, Alex Warstadt, Yohei Oseki, Ethan Gotlieb WilcoxACL 2026 · 1 citation
