ACL2026

Lending Eyesight to Language Models: Modeling and Probing Human scanpath through Transformer Decoder

Junlin Li, David Robert Reich, Yu-Yin Hsu

Abstract

Human scanpaths offer rich and reliable clues about the cognitive mechanisms underlying language comprehension. Decoder-only language models, typically large language models (LLMs), have proven to exhibit striking parallels with human cognitive processes. In this study, we investigate to what extent language models can be endowed with humanlike gaze shifts. Besides, by probing scanpath through eye model, analogous to probing language through language models, we ask whether such modeling can yield novel knowledge of the cognitive machinery of sense making. This study presents a novel "plug-in" module, EyeLM, to transform an autoregressive language model into an autoregressive eye model, thus facilitating a probabilistic spatial modeling of human explicit attention. Our EyeLM module, powered by LLMs, achieves competitive performance with novel cognitive probing capabilities. By probing EyeLM, we can reach the predictability and uncertainty of the scanpath. Exhibiting aligned patterns with prior knowledge about human reading comprehension, these probabilistic measures of scanpath act as promising predictors of human comprehension skills. 1