Human Attention in Image Captioning: Dataset and Analysis
Sen He, Hamed Rezazadegan Tavakoli, Ali Borji, Nicolas Pugeault
摘要
In this work, we present a novel dataset consisting of eye movements and verbal descriptions recorded synchronously over images. Using this data, we study the differences in human attention during free-viewing and image captioning tasks. We look into the relationship between human atten- tion and language constructs during perception and sen- tence articulation. We also analyse attention deployment mechanisms in the top-down soft attention approach that is argued to mimic human attention in captioning tasks, and investigate whether visual saliency can help image caption- ing. Our study reveals that (1) human attention behaviour differs in free-viewing and image description tasks. Hu- mans tend to fixate on a greater variety of regions under the latter task, (2) there is a strong relationship between de- scribed objects and attended objects (97% of the described objects are being attended), (3) a convolutional neural net- work as feature encoder accounts for human-attended re- gions during image captioning to a great extent (around 78%), (4) soft-attention mechanism differs from human at- tention, both spatially and temporally, and there is low correlation between caption scores and attention consis- tency scores. These indicate a large gap between humans and machines in regards to top-down attention, and (5) by integrating the soft attention model with image saliency, we can significantly improve the model’s performance on Flickr30k and MSCOCO benchmarks. The dataset can be found at: https://github.com/SenHe/ Human-Attention-in-Image-Captioning.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper7
- Machine versus Human Attention in Deep Reinforcement Learning TasksSihang Guo, Ruohan Zhang, Bo Liu, Yifeng Zhu 等NeurIPS 2021 · 被引用 38 次
- The Value of AI Guidance in Human Examination of Synthetically-Generated FacesAidan Boyd, Patrick Tinsley, Kevin W. Bowyer, Adam CzajkaAAAI 2023 · 被引用 20 次
- Topic Scene Graph Generation by Attention Distillation from CaptionWenbin Wang, Ruiping Wang, Xilin ChenICCV 2021 · 被引用 16 次
- Modeling Human Gaze Behavior with Diffusion Models for Unified Scanpath PredictionGiuseppe Cartella, Vittorio Cuculo, Alessandro D'Amelio, Marcella Cornia 等ICCV 2025 · 被引用 3 次
- Personalized Image Descriptions from Attention SequencesRuoyu Xue, Hieu Le, Jingyi Xu, Sounak Mondal 等CVPR 2026 · 被引用 2 次
相关 Paper
- Exploring Language Prior for Mode-Sensitive Visual Attention ModelingXiaoshuai Sun, Xuying Zhang, Liujuan Cao, Yongjian Wu 等ACM MM 2020 · 被引用 3 次
- Generating Image Descriptions via Sequential Cross-Modal Alignment Guided by Human GazeEce Takmaz, Sandro Pezzelle, Lisa Beinborn, Raquel FernándezEMNLP 2020 · 被引用 1 次
- Saliency-Guided Attention Network for Image-Sentence MatchingZhong Ji, Haoran Wang, Jungong Han, Yanwei PangICCV 2019 · 被引用 96 次
- Salient Object Ranking via Cyclical Perception-Viewing Interaction ModelingRongjin Guo, Ke Xu, Rynson W. H. LauICLR 2026
- Improving Image Captioning through Visual and Semantic Mutual PromotionJing Zhang, Yingshuai Xie, Xiaoqiang LiuACM MM 2023 · 被引用 4 次
