Cross-modal Coherence Modeling for Caption Generation
Malihe Alikhani, Piyush Sharma, Shengjie Li, Radu Soricut, Matthew Stone
摘要
We use coherence relations inspired by computational models of discourse to study the information needs and goals of image captioning. Using an annotation protocol specifically devised for capturing image-caption coherence relations, we annotate 10,000 instances from publicly-available image-caption pairs. We introduce a new task for learning inferences in imagery and text, coherence relation prediction, and show that these coherence annotations can be exploited to learn relation classifiers as an intermediary step, and also train coherence-aware, controllable image captioning models. The results show a dramatic improvement in the consistency and quality of the generated captions with respect to information needs specified via coherence relations.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper9
- Crossmodal-3600: A Massively Multilingual Multimodal Evaluation DatasetAshish V. Thapliyal, Jordi Pont-Tuset, Xi Chen, Radu SoricutEMNLP 2022 · 被引用 31 次
- AESOP: Abstract Encoding of Stories, Objects, and PicturesHareesh Ravi, Kushal Kafle, Scott Cohen, Jonathan Brandt 等ICCV 2021 · 被引用 19 次
- Top-Down Semantic Refinement for Image CaptioningJusheng Zhang, Kaitong Cai, Jing Yang, Jian Wang 等AAAI 2026 · 被引用 16 次
- MemeCap: A Dataset for Captioning and Interpreting MemesEunjeong Hwang, Vered ShwartzEMNLP 2023 · 被引用 14 次
- Concadia: Towards Image-Based Text Generation with a PurposeElisa Kreiss, Fei Fang, Noah D. Goodman, Christopher PottsEMNLP 2022 · 被引用 13 次
相关 Paper
- Cross-Modal Coherence for Text-to-Image RetrievalMalihe Alikhani, Fangda Han, Hareesh Ravi, Mubbasir Kapadia 等AAAI 2022 · 被引用 11 次
- CapWAP: Image Captioning with a PurposeAdam Fisch, Kenton Lee, Ming-Wei Chang, Jonathan H. Clark 等EMNLP 2020 · 被引用 17 次
- Relational Distant Supervision for Image Captioning without Image-Text PairsYayun Qi, Wentian Zhao, Xinxiao WuAAAI 2024 · 被引用 5 次
- Reflective Decoding Network for Image CaptioningLei Ke, Wenjie Pei, Ruiyu Li, Xiaoyong Shen 等ICCV 2019 · 被引用 107 次
- MAGIC: Multimodal relAtional Graph adversarIal inferenCe for Diverse and Unpaired Text-Based Image CaptioningWenqiao Zhang, Haochen Shi, Jiannan Guo, Shengyu Zhang 等AAAI 2022 · 被引用 52 次
