Towards Unsupervised Image Captioning With Shared Multimodal Embeddings
Iro Laina, Christian Rupprecht, Nassir Navab
Abstract
Understanding images without explicit supervision has become an important problem in computer vision. In this paper, we address image captioning by generating language descriptions of scenes without learning from annotated pairs of images and their captions. The core component of our approach is a shared latent space that is structured by visual concepts. In this space, the two modalities should be indistinguishable. A language model is first trained to encode sentences into semantically structured embeddings. Image features that are translated into this embedding space can be decoded into descriptions through the same language model, similarly to sentence embeddings. This translation is learned from weakly paired images and text using a loss robust to noisy assignments and a conditional adversarial component. Our approach allows to exploit large text corpora outside the annotated distributions of image/caption data. Our experiments show that the proposed domain alignment learns a semantically meaningful representation which outperforms previous work.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext dc56257b-b262-4e97-9b17-fd52b41b5750Cited by top-tier papers21
- ZeroCap: Zero-Shot Image-to-Text Generation for Visual-Semantic ArithmeticYoad Tewel, Yoav Shalev, Idan Schwartz, Lior WolfCVPR 2022 · 129 citations
- Transferable Decoding with Visual Entities for Zero-Shot Image CaptioningJunjie Fei, Teng Wang, Jinrui Zhang, Zhenyu He et al.ICCV 2023 · 80 citations
- Data Poisoning Attacks Against Multimodal EncodersZiqing Yang, Xinlei He, Zheng Li, Michael Backes et al.ICML 2023 · 74 citations
- Zero-shot Natural Language Video LocalizationJinwoo Nam, Daechul Ahn, Dongyeop Kang, Seong Jong Ha et al.ICCV 2021 · 60 citations
- Pseudo-Q: Generating Pseudo Language Queries for Visual GroundingHaojun Jiang, Yuanze Lin, Dongchen Han, Shiji Song et al.CVPR 2022 · 60 citations
Related papers
- Unpaired Image Captioning via Scene Graph AlignmentsJiuxiang Gu, Shafiq R. Joty, Jianfei Cai, Handong Zhao et al.ICCV 2019 · 191 citations
- Detach and Attach: Stylized Image Captioning without Paired Stylized DatasetYutong Tan, Zheng Lin, Peng Fu, Mingyu Zheng et al.ACM MM 2022 · 8 citations
- UNISON: Unpaired Cross-Lingual Image CaptioningJiahui Gao, Yi Zhou, Philip L. H. Yu, Shafiq R. Joty et al.AAAI 2022 · 18 citations
- Relational Distant Supervision for Image Captioning without Image-Text PairsYayun Qi, Wentian Zhao, Xinxiao WuAAAI 2024 · 5 citations
- Latent Normalizing Flows for Many-to-Many Cross-Domain MappingsShweta Mahajan, Iryna Gurevych, Stefan RothICLR 2020 · 38 citations
