Generating Image Descriptions via Sequential Cross-Modal Alignment Guided by Human Gaze
Ece Takmaz, Sandro Pezzelle, Lisa Beinborn, Raquel Fernández
Abstract
When speakers describe an image, they tend to look at objects before mentioning them. In this paper, we investigate such sequential crossmodal alignment by modelling the image description generation process computationally. We take as our starting point a state-of-theart image captioning system and develop several model variants that exploit information from human gaze patterns recorded during language production. In particular, we propose the first approach to image description generation where visual processing is modelled sequentially. Our experiments and analyses confirm that better descriptions can be obtained by exploiting gaze-driven attention and shed light on human cognitive processes by comparing different ways of aligning the gaze modality with language production. We find that processing gaze data sequentially leads to descriptions that are better aligned to those produced by speakers, more diverse, and more naturalparticularly when gaze is encoded with a dedicated recurrent component.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 6c144a6d-4906-44b0-8103-80a1d13beb83Cited by top-tier papers3
- Neglected Free Lunch - Learning Image Classifiers Using Annotation ByproductsDongyoon Han, Junsuk Choe, Seonghyeok Chun, John Joon Young Chung et al.ICCV 2023 · 4 citations
- Personalized Image Descriptions from Attention SequencesRuoyu Xue, Hieu Le, Jingyi Xu, Sounak Mondal et al.CVPR 2026 · 2 citations
- CogAlign: Learning to Align Textual Neural Representations to Cognitive Language Processing SignalsYuqi Ren, Deyi XiongACL 2021
Builds on1
Related papers
- CARIS: Context-Aware Referring Image SegmentationSun'ao Liu, Yiheng Zhang, Zhaofan Qiu, Hongtao Xie et al.ACM MM 2023 · 34 citations
- Iterative Back Modification for Faster Image CaptioningZhengcong FeiACM MM 2020 · 27 citations
- Exploring Overall Contextual Information for Image Captioning in Human-Like Cognitive StyleHongwei Ge, Zehang Yan, Kai Zhang, Mingde Zhao et al.ICCV 2019 · 25 citations
- Towards Unsupervised Image Captioning With Shared Multimodal EmbeddingsIro Laina, Christian Rupprecht, Nassir NavabICCV 2019 · 115 citations
- Reflective Decoding Network for Image CaptioningLei Ke, Wenjie Pei, Ruiyu Li, Xiaoyong Shen et al.ICCV 2019 · 107 citations
