Predicting Goal-Directed Human Attention Using Inverse Reinforcement Learning
Zhibo Yang, Lihan Huang, Yupei Chen, Zijun Wei, Seoyoung Ahn, Gregory J. Zelinsky, Dimitris Samaras, Minh Hoai
摘要
Being able to predict human gaze behavior has obvious importance for behavioral vision and for computer vision applications. Most models have mainly focused on predicting free-viewing behavior using saliency maps, but these predictions do not generalize to goal-directed behavior, such as when a person searches for a visual target object. We propose the first inverse reinforcement learning (IRL) model to learn the internal reward function and policy used by humans during visual search. The viewer's internal belief states were modeled as dynamic contextual belief maps of object locations. These maps were learned by IRL and then used to predict behavioral scanpaths for multiple target categories. To train and evaluate our IRL model we created COCO-Search18, which is now the largest dataset of high-quality search fixations in existence. COCO-Search18 has 10 participants searching for each of 18 target-object categories in 6202 images, making about 300,000 goaldirected fixations. When trained and evaluated on COCO-Search18, the IRL model outperformed baseline models in predicting search fixation scanpaths, both in terms of similarity to human search behavior and search efficiency. Finally, reward maps recovered by the IRL model reveal distinctive target-dependent patterns of object prioritization, which we interpret as a learned object context. Code and dataset are available at https://github.com/ cvlab-stonybrook/Scanpath_Prediction . *Equal contribution.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper27
- MEDIRL: Predicting the Visual Attention of Drivers via Maximum Entropy Deep Inverse Reinforcement LearningSonia Baee, Erfan Pakdamanian, Inki Kim, Lu Feng 等ICCV 2021 · 被引用 65 次
- DRIVE: Deep Reinforced Accident Anticipation with Visual ExplanationWentao Bao, Qi Yu, Yu KongICCV 2021 · 被引用 64 次
- V*: Guided Visual Search as a Core Mechanism in Multimodal LLMsPenghao Wu, Saining XieCVPR 2024 · 被引用 32 次
- UniAR: A Unified model for predicting human Attention and Responses on visual contentPeizhao Li, Junfeng He, Gang Li, Rachit Bhargava 等NeurIPS 2024 · 被引用 17 次
- SalChartQA: Question-driven Saliency on Information VisualisationsYao Wang, Weitian Wang, Abdullah Abdelhafez, Mayar Elfares 等CHI 2024 · 被引用 15 次
相关 Paper
- PRE-MAP: Personalized Reinforced Eye-tracking Multimodal LLM for High-Resolution Multi-Attribute Point PredictionHanbing Wu, Ping Jiang, Anyang Su, Chenxu Zhao 等ACM MM 2025
- Predicting Human Scanpaths in Visual Question AnsweringXianyu Chen, Ming Jiang, Qi ZhaoCVPR 2021
- SeekUI: Predicting Visual Search Behavior on Graphical User Interfaces with a Reward-Augmented Vision Language ModelZixin Guo, Yue Jiang, Luis A. Leiva, Antti OulasvirtaCHI 2026 · 被引用 1 次
- Dynamic Inverse Reinforcement Learning for Characterizing Animal BehaviorZoe Ashwood, Aditi Jha, Jonathan W. PillowNeurIPS 2022 · 被引用 50 次
- EyeFormer: Predicting Personalized Scanpaths with Transformer-Guided Reinforcement LearningYue Jiang, Zixin Guo, Hamed Rezazadegan Tavakoli, Luis A. Leiva 等UIST 2024 · 被引用 15 次
