Predicting Human Scanpaths in Visual Question Answering
Xianyu Chen, Ming Jiang, Qi Zhao
Abstract
Attention has been an important mechanism for both humans and computer vision systems. While state-of-theart models to predict attention focus on estimating a static probabilistic saliency map with free-viewing behavior, reallife scenarios are filled with tasks of varying types and complexities, and visual exploration is a temporal process that contributes to task performance. To bridge the gap, we conduct a first study to understand and predict the temporal sequences of eye fixations (a.k.a. scanpaths) during performing general tasks, and examine how scanpaths affect task performance. We present a new deep reinforcement learning method to predict scanpaths leading to different performances in visual question answering. Conditioned on a task guidance map, the proposed model learns question-specific attention patterns to generate scanpaths. It addresses the exposure bias in scanpath prediction with self-critical sequence training and designs a Consistency-Divergence loss to generate distinguishable scanpaths between correct and incorrect answers. The proposed model not only accurately predicts the spatio-temporal patterns of human behavior in visual question answering, such as fixation position, duration, and order, but also generalizes to free-viewing and visual search tasks, achieving human-level performance in all tasks and significantly outperforming the state of the art. Question: Is the vase the same color as the scarf?
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b98438fe-b963-4212-8e8f-77b54b8b954aCited by top-tier papers21
- UniAR: A Unified model for predicting human Attention and Responses on visual contentPeizhao Li, Junfeng He, Gang Li, Rachit Bhargava et al.NeurIPS 2024 · 17 citations
- EyeFormer: Predicting Personalized Scanpaths with Transformer-Guided Reinforcement LearningYue Jiang, Zixin Guo, Hamed Rezazadegan Tavakoli, Luis A. Leiva et al.UIST 2024 · 15 citations
- Chartist: Task-driven Eye Movement Control for Chart ReadingDanqing Shi, Yao Wang, Yunpeng Bai, Andreas Bulling et al.CHI 2025 · 13 citations
- ScanTD: 360° Scanpath Prediction based on Time-Series DiffusionYujia Wang, Fang-Lue Zhang, Neil A. DodgsonACM MM 2024 · 11 citations
- DiffEye: Diffusion-Based Continuous Eye-Tracking Data Generation Conditioned on Natural ImagesOzgur Kara, Harris Nisar, James M. RehgNeurIPS 2025 · 7 citations
Builds on2
Related papers
- Beyond Average: Individualized Visual Scanpath PredictionXianyu Chen, Ming Jiang, Qi ZhaoCVPR 2024
- Modeling Human Gaze Behavior with Diffusion Models for Unified Scanpath PredictionGiuseppe Cartella, Vittorio Cuculo, Alessandro D'Amelio, Marcella Cornia et al.ICCV 2025 · 3 citations
- PRE-MAP: Personalized Reinforced Eye-tracking Multimodal LLM for High-Resolution Multi-Attribute Point PredictionHanbing Wu, Ping Jiang, Anyang Su, Chenxu Zhao et al.ACM MM 2025
- What Moves the Eyes: Doubling Mechanistic Model Performance Using Deep Networks to Discover and Test Cognitive HypothesesFederico D'Agostino, Lisa Schwetlick, Matthias Bethge, Matthias KümmererNeurIPS 2025 · 4 citations
- ScanDMM: A Deep Markov Model of Scanpath Prediction for 360° ImagesXiangjie Sui, Yuming Fang, Hanwei Zhu, Shiqi Wang et al.CVPR 2023
