MAGIC: Multimodal relAtional Graph adversarIal inferenCe for Diverse and Unpaired Text-Based Image Captioning
Wenqiao Zhang, Haochen Shi, Jiannan Guo, Shengyu Zhang, Qingpeng Cai, Juncheng Li, Sihui Luo, Yueting Zhuang
Abstract
Text-based image captioning (TextCap) requires simultaneous comprehension of visual content and reading the text of images to generate a natural language description. Although a task can teach machines to understand the complex human environment further given that text is omnipresent in our daily surroundings, it poses additional challenges in normal captioning. A text-based image intuitively contains abundant and complex multimodal relational content, that is, image details can be described diversely from multiview rather than a single caption. Certainly, we can introduce additional paired training data to show the diversity of images' descriptions, this process is labor-intensive and time-consuming for TextCap pair annotations with extra texts. Based on the insight mentioned above, we investigate how to generate diverse captions that focus on different image parts using an unpaired training paradigm. We propose the Multimodal relAtional Graph adversarIal InferenCe (MAGIC) framework for diverse and unpaired TextCap. This framework can adaptively construct multiple multimodal relational graphs of images and model complex relationships among graphs to represent descriptive diversity. Moreover, a cascaded generative adversarial network is developed from modeled graphs to infer the unpaired caption generation in image–sentence feature alignment and linguistic coherence levels. We validate the effectiveness of MAGIC in generating diverse captions from different relational information items of an image. Experimental results show that MAGIC can generate very promising outcomes without using any image–caption training pairs.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 829a7fa1-a586-45d2-bb29-ab665d4a272bCited by top-tier papers21
- BoostMIS: Boosting Medical Image Semi-supervised Learning with Adaptive Pseudo Labeling and Informative Active AnnotationWenqiao Zhang, Lei Zhu, James Hallinan, Shengyu Zhang et al.CVPR 2022 · 115 citations
- Fine-tuning Multimodal LLMs to Follow Zero-shot Demonstrative InstructionsJuncheng Li, Kaihang Pan, Zhiqi Ge, Minghe Gao et al.ICLR 2024 · 95 citations
- Re4: Learning to Re-contrast, Re-attend, Re-construct for Multi-interest RecommendationShengyu Zhang, Lingxiao Yang, Dong Yao, Yujie Lu et al.WWW 2022 · 67 citations
- Compositional Temporal Grounding with Structured Variational Cross-Graph Correspondence LearningJuncheng Li, Junlin Xie, Long Qian, Linchao Zhu et al.CVPR 2022 · 63 citations
- Intelligent Model Update Strategy for Sequential RecommendationZheqi Lv, Wenqiao Zhang, Zhengyu Chen, Shengyu Zhang et al.WWW 2024 · 53 citations
Builds on9
- Unsupervised Learning of Visual Features by Contrasting Cluster AssignmentsMathilde Caron, Ishan Misra, Julien Mairal, Priya Goyal et al.NeurIPS 2020 · 5,249 citations
- Unpaired Image Captioning via Scene Graph AlignmentsJiuxiang Gu, Shafiq R. Joty, Jianfei Cai, Handong Zhao et al.ICCV 2019 · 191 citations
- Semi-supervised Active Learning for Semi-supervised Models: Exploit Adversarial Examples with Graph-based Virtual LabelsJiannan Guo, Haochen Shi, Yangyang Kang, Kun Kuang et al.ICCV 2021 · 38 citations
- Adaptive Hierarchical Graph Reasoning with Semantic Coherence for Video-and-Language InferenceJuncheng Li, Siliang Tang, Linchao Zhu, Haochen Shi et al.ICCV 2021 · 28 citations
- Relational Graph Learning for Grounded Video Description GenerationWenqiao Zhang, Xin Eric Wang, Siliang Tang, Haizhou Shi et al.ACM MM 2020 · 25 citations
Related papers
- Towards Accurate Text-Based Image Captioning With Content Diversity ExplorationGuanghui Xu, Shuaicheng Niu, Mingkui Tan, Yucheng Luo et al.CVPR 2021
- UNISON: Unpaired Cross-Lingual Image CaptioningJiahui Gao, Yi Zhou, Philip L. H. Yu, Shafiq R. Joty et al.AAAI 2022 · 18 citations
- Interactive Dual Generative Adversarial Networks for Image CaptioningJunhao Liu, Kai Wang, Chunpu Xu, Zhou Zhao et al.AAAI 2020 · 35 citations
- Noise-Aware Image Captioning with Progressively Exploring Mismatched WordsZhongtian Fu, Kefei Song, Luping Zhou, Yang YangAAAI 2024 · 36 citations
- CapOnImage: Context-driven Dense-Captioning on ImageYiqi Gao, Xinglin Hou, Yuanmeng Zhang, Tiezheng Ge et al.EMNLP 2022 · 5 citations
