Relational Distant Supervision for Image Captioning without Image-Text Pairs
Yayun Qi, Wentian Zhao, Xinxiao Wu
Abstract
Unsupervised image captioning aims to generate descriptions of images without relying on any image-sentence pairs for training. Most existing works use detected visual objects or concepts as bridge to connect images and texts. Considering that the relationship between objects carries more information, we use the object relationship as a more accurate connection between images and texts. In this paper, we adapt the idea of distant supervision that extracts the knowledge about object relationships from an external corpus and imparts them to images to facilitate inferring visual object relationships, without introducing any extra pre-trained relationship detectors. Based on these learned informative relationships, we construct pseudo image-sentence pairs for captioning model training. Specifically, our method consists of three modules: (i) a relationship learning module that learns to infer relationships from images under the distant supervision; (ii) a relationship-to-sentence module that transforms the inferred relationships into sentences to generate pseudo image-sentence pairs; (iii) an image captioning module that is trained by using the generated image-sentence pairs. Promising results on three datasets show that our method outperforms the state-of-the-art methods of unsupervised image captioning.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on5
- BOND: BERT-Assisted Open-Domain Named Entity Recognition with Distant SupervisionChen Liang, Yue Yu, Haoming Jiang, Siawpeng Er et al.KDD 2020 · 118 citations
- Towards Unsupervised Image Captioning With Shared Multimodal EmbeddingsIro Laina, Christian Rupprecht, Nassir NavabICCV 2019 · 115 citations
- Distantly-Supervised Named Entity Recognition with Noise-Robust Learning and Language Model Augmented Self-TrainingYu Meng, Yunyi Zhang, Jiaxin Huang, Xuan Wang et al.EMNLP 2021 · 50 citations
- Visual Distant Supervision for Scene Graph GenerationYuan Yao, Ao Zhang, Xu Han, Mengdi Li et al.ICCV 2021 · 41 citations
- Unbiased Scene Graph Generation From Biased TrainingKaihua Tang, Yulei Niu, Jianqiang Huang, Jiaxin Shi et al.CVPR 2020
Related papers
- Joint Commonsense and Relation Reasoning for Image and Video CaptioningJingyi Hou, Xinxiao Wu, Xiaoxun Zhang, Yayun Qi et al.AAAI 2020 · 52 citations
- Bridging the Gap between Vision and Language Domains for Improved Image CaptioningFenglin Liu, Xian Wu, Shen Ge, Xiaoyu Zhang et al.ACM MM 2020 · 13 citations
- Learning to Generate Scene Graph from Natural Language SupervisionYiwu Zhong, Jing Shi, Jianwei Yang, Chenliang Xu et al.ICCV 2021 · 88 citations
- Learning Human-Human Interactions in Images from Weak Textual SupervisionMorris Alper, Hadar Averbuch-ElorICCV 2023 · 4 citations
- Are Noisy Sentences Useless for Distant Supervised Relation Extraction?Yuming Shang, He Yan Huang, Xianling Mao, Xin Sun et al.AAAI 2020 · 39 citations
