Cross-modal Coherence Modeling for Caption Generation
Malihe Alikhani, Piyush Sharma, Shengjie Li, Radu Soricut, Matthew Stone
Abstract
We use coherence relations inspired by computational models of discourse to study the information needs and goals of image captioning. Using an annotation protocol specifically devised for capturing image-caption coherence relations, we annotate 10,000 instances from publicly-available image-caption pairs. We introduce a new task for learning inferences in imagery and text, coherence relation prediction, and show that these coherence annotations can be exploited to learn relation classifiers as an intermediary step, and also train coherence-aware, controllable image captioning models. The results show a dramatic improvement in the consistency and quality of the generated captions with respect to information needs specified via coherence relations.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b6a1269f-533a-4f3f-a724-71ed364f77faCited by top-tier papers9
- Crossmodal-3600: A Massively Multilingual Multimodal Evaluation DatasetAshish V. Thapliyal, Jordi Pont-Tuset, Xi Chen, Radu SoricutEMNLP 2022 · 31 citations
- AESOP: Abstract Encoding of Stories, Objects, and PicturesHareesh Ravi, Kushal Kafle, Scott Cohen, Jonathan Brandt et al.ICCV 2021 · 19 citations
- Top-Down Semantic Refinement for Image CaptioningJusheng Zhang, Kaitong Cai, Jing Yang, Jian Wang et al.AAAI 2026 · 16 citations
- MemeCap: A Dataset for Captioning and Interpreting MemesEunjeong Hwang, Vered ShwartzEMNLP 2023 · 14 citations
- Concadia: Towards Image-Based Text Generation with a PurposeElisa Kreiss, Fei Fang, Noah D. Goodman, Christopher PottsEMNLP 2022 · 13 citations
Related papers
- Cross-Modal Coherence for Text-to-Image RetrievalMalihe Alikhani, Fangda Han, Hareesh Ravi, Mubbasir Kapadia et al.AAAI 2022 · 11 citations
- CapWAP: Image Captioning with a PurposeAdam Fisch, Kenton Lee, Ming-Wei Chang, Jonathan H. Clark et al.EMNLP 2020 · 17 citations
- Relational Distant Supervision for Image Captioning without Image-Text PairsYayun Qi, Wentian Zhao, Xinxiao WuAAAI 2024 · 5 citations
- Reflective Decoding Network for Image CaptioningLei Ke, Wenjie Pei, Ruiyu Li, Xiaoyong Shen et al.ICCV 2019 · 107 citations
- MAGIC: Multimodal relAtional Graph adversarIal inferenCe for Diverse and Unpaired Text-Based Image CaptioningWenqiao Zhang, Haochen Shi, Jiannan Guo, Shengyu Zhang et al.AAAI 2022 · 52 citations
