MAGIC: Multimodal relAtional Graph adversarIal inferenCe for Diverse and Unpaired Text-Based Image Captioning
Wenqiao Zhang, Haochen Shi, Jiannan Guo, Shengyu Zhang, Qingpeng Cai, Juncheng Li, Sihui Luo, Yueting Zhuang
摘要
Text-based image captioning (TextCap) requires simultaneous comprehension of visual content and reading the text of images to generate a natural language description. Although a task can teach machines to understand the complex human environment further given that text is omnipresent in our daily surroundings, it poses additional challenges in normal captioning. A text-based image intuitively contains abundant and complex multimodal relational content, that is, image details can be described diversely from multiview rather than a single caption. Certainly, we can introduce additional paired training data to show the diversity of images' descriptions, this process is labor-intensive and time-consuming for TextCap pair annotations with extra texts. Based on the insight mentioned above, we investigate how to generate diverse captions that focus on different image parts using an unpaired training paradigm. We propose the Multimodal relAtional Graph adversarIal InferenCe (MAGIC) framework for diverse and unpaired TextCap. This framework can adaptively construct multiple multimodal relational graphs of images and model complex relationships among graphs to represent descriptive diversity. Moreover, a cascaded generative adversarial network is developed from modeled graphs to infer the unpaired caption generation in image–sentence feature alignment and linguistic coherence levels. We validate the effectiveness of MAGIC in generating diverse captions from different relational information items of an image. Experimental results show that MAGIC can generate very promising outcomes without using any image–caption training pairs.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper21
- BoostMIS: Boosting Medical Image Semi-supervised Learning with Adaptive Pseudo Labeling and Informative Active AnnotationWenqiao Zhang, Lei Zhu, James Hallinan, Shengyu Zhang 等CVPR 2022 · 被引用 115 次
- Fine-tuning Multimodal LLMs to Follow Zero-shot Demonstrative InstructionsJuncheng Li, Kaihang Pan, Zhiqi Ge, Minghe Gao 等ICLR 2024 · 被引用 95 次
- Re4: Learning to Re-contrast, Re-attend, Re-construct for Multi-interest RecommendationShengyu Zhang, Lingxiao Yang, Dong Yao, Yujie Lu 等WWW 2022 · 被引用 67 次
- Compositional Temporal Grounding with Structured Variational Cross-Graph Correspondence LearningJuncheng Li, Junlin Xie, Long Qian, Linchao Zhu 等CVPR 2022 · 被引用 63 次
- Intelligent Model Update Strategy for Sequential RecommendationZheqi Lv, Wenqiao Zhang, Zhengyu Chen, Shengyu Zhang 等WWW 2024 · 被引用 53 次
它引用的顶会 Paper9
- Unsupervised Learning of Visual Features by Contrasting Cluster AssignmentsMathilde Caron, Ishan Misra, Julien Mairal, Priya Goyal 等NeurIPS 2020 · 被引用 5,249 次
- Unpaired Image Captioning via Scene Graph AlignmentsJiuxiang Gu, Shafiq R. Joty, Jianfei Cai, Handong Zhao 等ICCV 2019 · 被引用 191 次
- Semi-supervised Active Learning for Semi-supervised Models: Exploit Adversarial Examples with Graph-based Virtual LabelsJiannan Guo, Haochen Shi, Yangyang Kang, Kun Kuang 等ICCV 2021 · 被引用 38 次
- Adaptive Hierarchical Graph Reasoning with Semantic Coherence for Video-and-Language InferenceJuncheng Li, Siliang Tang, Linchao Zhu, Haochen Shi 等ICCV 2021 · 被引用 28 次
- Relational Graph Learning for Grounded Video Description GenerationWenqiao Zhang, Xin Eric Wang, Siliang Tang, Haizhou Shi 等ACM MM 2020 · 被引用 25 次
相关 Paper
- Towards Accurate Text-Based Image Captioning With Content Diversity ExplorationGuanghui Xu, Shuaicheng Niu, Mingkui Tan, Yucheng Luo 等CVPR 2021
- UNISON: Unpaired Cross-Lingual Image CaptioningJiahui Gao, Yi Zhou, Philip L. H. Yu, Shafiq R. Joty 等AAAI 2022 · 被引用 18 次
- Interactive Dual Generative Adversarial Networks for Image CaptioningJunhao Liu, Kai Wang, Chunpu Xu, Zhou Zhao 等AAAI 2020 · 被引用 35 次
- Noise-Aware Image Captioning with Progressively Exploring Mismatched WordsZhongtian Fu, Kefei Song, Luping Zhou, Yang YangAAAI 2024 · 被引用 36 次
- CapOnImage: Context-driven Dense-Captioning on ImageYiqi Gao, Xinglin Hou, Yuanmeng Zhang, Tiezheng Ge 等EMNLP 2022 · 被引用 5 次
