Predicting Human Scanpaths in Visual Question Answering
Xianyu Chen, Ming Jiang, Qi Zhao
摘要
Attention has been an important mechanism for both humans and computer vision systems. While state-of-theart models to predict attention focus on estimating a static probabilistic saliency map with free-viewing behavior, reallife scenarios are filled with tasks of varying types and complexities, and visual exploration is a temporal process that contributes to task performance. To bridge the gap, we conduct a first study to understand and predict the temporal sequences of eye fixations (a.k.a. scanpaths) during performing general tasks, and examine how scanpaths affect task performance. We present a new deep reinforcement learning method to predict scanpaths leading to different performances in visual question answering. Conditioned on a task guidance map, the proposed model learns question-specific attention patterns to generate scanpaths. It addresses the exposure bias in scanpath prediction with self-critical sequence training and designs a Consistency-Divergence loss to generate distinguishable scanpaths between correct and incorrect answers. The proposed model not only accurately predicts the spatio-temporal patterns of human behavior in visual question answering, such as fixation position, duration, and order, but also generalizes to free-viewing and visual search tasks, achieving human-level performance in all tasks and significantly outperforming the state of the art. Question: Is the vase the same color as the scarf?
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper21
- UniAR: A Unified model for predicting human Attention and Responses on visual contentPeizhao Li, Junfeng He, Gang Li, Rachit Bhargava 等NeurIPS 2024 · 被引用 17 次
- EyeFormer: Predicting Personalized Scanpaths with Transformer-Guided Reinforcement LearningYue Jiang, Zixin Guo, Hamed Rezazadegan Tavakoli, Luis A. Leiva 等UIST 2024 · 被引用 15 次
- Chartist: Task-driven Eye Movement Control for Chart ReadingDanqing Shi, Yao Wang, Yunpeng Bai, Andreas Bulling 等CHI 2025 · 被引用 13 次
- ScanTD: 360° Scanpath Prediction based on Time-Series DiffusionYujia Wang, Fang-Lue Zhang, Neil A. DodgsonACM MM 2024 · 被引用 11 次
- DiffEye: Diffusion-Based Continuous Eye-Tracking Data Generation Conditioned on Natural ImagesOzgur Kara, Harris Nisar, James M. RehgNeurIPS 2025 · 被引用 7 次
它引用的顶会 Paper2
相关 Paper
- Beyond Average: Individualized Visual Scanpath PredictionXianyu Chen, Ming Jiang, Qi ZhaoCVPR 2024
- Modeling Human Gaze Behavior with Diffusion Models for Unified Scanpath PredictionGiuseppe Cartella, Vittorio Cuculo, Alessandro D'Amelio, Marcella Cornia 等ICCV 2025 · 被引用 3 次
- PRE-MAP: Personalized Reinforced Eye-tracking Multimodal LLM for High-Resolution Multi-Attribute Point PredictionHanbing Wu, Ping Jiang, Anyang Su, Chenxu Zhao 等ACM MM 2025
- What Moves the Eyes: Doubling Mechanistic Model Performance Using Deep Networks to Discover and Test Cognitive HypothesesFederico D'Agostino, Lisa Schwetlick, Matthias Bethge, Matthias KümmererNeurIPS 2025 · 被引用 4 次
- ScanDMM: A Deep Markov Model of Scanpath Prediction for 360° ImagesXiangjie Sui, Yuming Fang, Hanwei Zhu, Shiqi Wang 等CVPR 2023
