EyeFormer: Predicting Personalized Scanpaths with Transformer-Guided Reinforcement Learning
Yue Jiang, Zixin Guo, Hamed Rezazadegan Tavakoli, Luis A. Leiva, Antti Oulasvirta
Abstract
From a visual-perception perspective, modern graphical user interfaces (GUIs) comprise a complex graphics-rich two-dimensional visuospatial arrangement of text, images, and interactive objects such as buttons and menus. While existing models can accurately predict regions and objects that are likely to attract attention “on average”, no scanpath model has been capable of predicting scanpaths for an individual. To close this gap, we introduce EyeFormer, which utilizes a Transformer architecture as a policy network to guide a deep reinforcement learning algorithm that predicts gaze locations. Our model offers the unique capability of producing personalized predictions when given a few user scanpath samples. It can predict full scanpath information, including fixation positions and durations, across individuals and various stimulus types. Additionally, we demonstrate applications in GUI layout optimization driven by our model.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers5
- Chartist: Task-driven Eye Movement Control for Chart ReadingDanqing Shi, Yao Wang, Yunpeng Bai, Andreas Bulling et al.CHI 2025 · 13 citations
- Branch Explorer: Leveraging Branching Narratives to Support Interactive 360° Video Viewing for Blind and Low Vision UsersShuchang Xu, Xiaofu Jin, Wenshuo Zhang, Huamin Qu et al.UIST 2025 · 4 citations
- What Moves the Eyes: Doubling Mechanistic Model Performance Using Deep Networks to Discover and Test Cognitive HypothesesFederico D'Agostino, Lisa Schwetlick, Matthias Bethge, Matthias KümmererNeurIPS 2025 · 4 citations
- Forecasting 3D Scanpaths in Egocentric VideoFiona Ryan, Ishwarya Ananthabhotla, Yijun Qian, Judy Hoffman et al.CVPR 2026 · 1 citation
- Few-shot Personalized Scanpath PredictionRuoyu Xue, Jingyi Xu, Sounak Mondal, Hieu Le et al.CVPR 2025
Builds on10
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- When Vision Transformers Outperform ResNets without Pre-training or Strong Data AugmentationsXiangning Chen, Cho-Jui Hsieh, Boqing GongICLR 2022 · 388 citations
- Discretizing Continuous Action Space for On-Policy OptimizationYunhao Tang, Shipra AgrawalAAAI 2020 · 150 citations
- UEyes: Understanding Visual Saliency across User Interface TypesYue Jiang, Luis A. Leiva, Hamed Rezazadegan Tavakoli, Paul R. B. Houssel et al.CHI 2023 · 100 citations
- ScanGAN360: A Generative Model of Realistic Scanpaths for 360° ImagesDaniel Martin, Ana Serrano, Alexander W. Bergman, Gordon Wetzstein et al.IEEE VR 2022 · 71 citations
Related papers
- SeekUI: Predicting Visual Search Behavior on Graphical User Interfaces with a Reward-Augmented Vision Language ModelZixin Guo, Yue Jiang, Luis A. Leiva, Antti OulasvirtaCHI 2026 · 1 citation
- Gazeformer: Scalable, Effective and Fast Prediction of Goal-Directed Human AttentionSounak Mondal, Zhibo Yang, Seoyoung Ahn, Dimitris Samaras et al.CVPR 2023
- Predicting Human Scanpaths in Visual Question AnsweringXianyu Chen, Ming Jiang, Qi ZhaoCVPR 2021
- SpFormer: Spatio-Temporal Modeling for Scanpaths with TransformerWenqi Zhong, Linzhi Yu, Chen Xia, Junwei Han et al.AAAI 2024 · 7 citations
- Modeling Human Gaze Behavior with Diffusion Models for Unified Scanpath PredictionGiuseppe Cartella, Vittorio Cuculo, Alessandro D'Amelio, Marcella Cornia et al.ICCV 2025 · 3 citations
