Generating Image Descriptions via Sequential Cross-Modal Alignment Guided by Human Gaze
Ece Takmaz, Sandro Pezzelle, Lisa Beinborn, Raquel Fernández
摘要
When speakers describe an image, they tend to look at objects before mentioning them. In this paper, we investigate such sequential crossmodal alignment by modelling the image description generation process computationally. We take as our starting point a state-of-theart image captioning system and develop several model variants that exploit information from human gaze patterns recorded during language production. In particular, we propose the first approach to image description generation where visual processing is modelled sequentially. Our experiments and analyses confirm that better descriptions can be obtained by exploiting gaze-driven attention and shed light on human cognitive processes by comparing different ways of aligning the gaze modality with language production. We find that processing gaze data sequentially leads to descriptions that are better aligned to those produced by speakers, more diverse, and more naturalparticularly when gaze is encoded with a dedicated recurrent component.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Neglected Free Lunch - Learning Image Classifiers Using Annotation ByproductsDongyoon Han, Junsuk Choe, Seonghyeok Chun, John Joon Young Chung 等ICCV 2023 · 被引用 4 次
- Personalized Image Descriptions from Attention SequencesRuoyu Xue, Hieu Le, Jingyi Xu, Sounak Mondal 等CVPR 2026 · 被引用 2 次
- CogAlign: Learning to Align Textual Neural Representations to Cognitive Language Processing SignalsYuqi Ren, Deyi XiongACL 2021
它引用的顶会 Paper1
相关 Paper
- CARIS: Context-Aware Referring Image SegmentationSun'ao Liu, Yiheng Zhang, Zhaofan Qiu, Hongtao Xie 等ACM MM 2023 · 被引用 34 次
- Iterative Back Modification for Faster Image CaptioningZhengcong FeiACM MM 2020 · 被引用 27 次
- Exploring Overall Contextual Information for Image Captioning in Human-Like Cognitive StyleHongwei Ge, Zehang Yan, Kai Zhang, Mingde Zhao 等ICCV 2019 · 被引用 25 次
- Towards Unsupervised Image Captioning With Shared Multimodal EmbeddingsIro Laina, Christian Rupprecht, Nassir NavabICCV 2019 · 被引用 115 次
- Reflective Decoding Network for Image CaptioningLei Ke, Wenjie Pei, Ruiyu Li, Xiaoyong Shen 等ICCV 2019 · 被引用 107 次
