Gazeformer: Scalable, Effective and Fast Prediction of Goal-Directed Human Attention
Sounak Mondal, Zhibo Yang, Seoyoung Ahn, Dimitris Samaras, Gregory J. Zelinsky, Minh Hoai
摘要
Predicting human gaze is important in Human-Computer Interaction (HCI). However, to practically serve HCI applications, gaze prediction models must be scalable, fast, and accurate in their spatial and temporal gaze predictions. Recent scanpath prediction models focus on goaldirected attention (search). Such models are limited in their application due to a common approach relying on trained target detectors for all possible objects, and the availability of human gaze data for their training (both not scalable). In response, we pose a new task called ZeroGaze, a new variant of zero-shot learning where gaze is predicted for never-before-searched objects, and we develop a novel model, Gazeformer, to solve the ZeroGaze problem. In contrast to existing methods using object detector modules, Gazeformer encodes the target using a natural language model, thus leveraging semantic similarities in scanpath prediction. We use a transformer-based encoderdecoder architecture because transformers are particularly useful for generating contextual representations. Gazeformer surpasses other models by a large margin (19%-70%) on the ZeroGaze setting. It also outperforms existing target-detection models on standard gaze prediction for both target-present and target-absent search tasks. In addition to its improved performance, Gazeformer is more than five times faster than the state-of-the-art target-present visual search model. Code can be found at https:// github.com/cvlab-stonybrook/Gazeformer/
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper18
- UniAR: A Unified model for predicting human Attention and Responses on visual contentPeizhao Li, Junfeng He, Gang Li, Rachit Bhargava 等NeurIPS 2024 · 被引用 17 次
- EyeFormer: Predicting Personalized Scanpaths with Transformer-Guided Reinforcement LearningYue Jiang, Zixin Guo, Hamed Rezazadegan Tavakoli, Luis A. Leiva 等UIST 2024 · 被引用 15 次
- Chartist: Task-driven Eye Movement Control for Chart ReadingDanqing Shi, Yao Wang, Yunpeng Bai, Andreas Bulling 等CHI 2025 · 被引用 13 次
- Learning from Observer Gaze: Zero-Shot Attention Prediction Oriented by Human-Object Interaction RecognitionYuchen Zhou, Linkai Liu, Chao GouCVPR 2024 · 被引用 13 次
- DiffEye: Diffusion-Based Continuous Eye-Tracking Data Generation Conditioned on Natural ImagesOzgur Kara, Harris Nisar, James M. RehgNeurIPS 2025 · 被引用 7 次
它引用的顶会 Paper7
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu 等ICCV 2021 · 被引用 31,683 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Training data-efficient image transformers & distillation through attentionHugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa 等ICML 2021 · 被引用 8,974 次
- MDETR - Modulated Detection for End-to-End Multi-Modal UnderstandingAishwarya Kamath, Mannat Singh, Yann LeCun, Gabriel Synnaeve 等ICCV 2021 · 被引用 1,114 次
- TubeDETR: Spatio-Temporal Video Grounding with TransformersAntoine Yang, Antoine Miech, Josef Sivic, Ivan Laptev 等CVPR 2022 · 被引用 87 次
相关 Paper
- Unifying Top-Down and Bottom-Up Scanpath Prediction Using TransformersZhibo Yang, Sounak Mondal, Seoyoung Ahn, Ruoyu Xue 等CVPR 2024
- Modeling Human Gaze Behavior with Diffusion Models for Unified Scanpath PredictionGiuseppe Cartella, Vittorio Cuculo, Alessandro D'Amelio, Marcella Cornia 等ICCV 2025 · 被引用 3 次
- Gaze-Language Alignment for Zero-Shot Prediction of Visual Search Targets from Human Gaze ScanpathsSounak Mondal, Naveen Sendhilnathan, Ting Zhang, Yue Liu 等ICCV 2025 · 被引用 2 次
- Goal-Oriented Gaze Estimation for Zero-Shot LearningYang Liu, Lei Zhou, Xiao Bai, Yifei Huang 等CVPR 2021
- Beyond Average: Individualized Visual Scanpath PredictionXianyu Chen, Ming Jiang, Qi ZhaoCVPR 2024
