UNISON: Unpaired Cross-Lingual Image Captioning
Jiahui Gao, Yi Zhou, Philip L. H. Yu, Shafiq R. Joty, Jiuxiang Gu
摘要
Image captioning has emerged as an interesting research field in recent years due to its broad application scenarios. The traditional paradigm of image captioning relies on paired image-caption datasets to train the model in a supervised manner. However, creating such paired datasets for every target language is prohibitively expensive, which hinders the extensibility of captioning technology and deprives a large part of the world population of its benefit. In this work, we present a novel unpaired cross-lingual method to generate image captions without relying on any caption corpus in the source or the target language. Specifically, our method consists of two phases: (1) a cross-lingual auto-encoding process, which utilizing a sentence parallel (bitext) corpus to learn the mapping from the source to the target language in the scene graph encoding space and decode sentences in the target language, and (2) a cross-modal unsupervised feature mapping, which seeks to map the encoded scene graph features from image modality to language modality. We verify the effectiveness of our proposed method on the Chinese image caption generation task. The comparisons against several existing methods demonstrate the effectiveness of our approach.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper8
- Sparse Invariant Risk MinimizationXiao Zhou, Yong Lin, Weizhong Zhang, Tong ZhangICML 2022 · 被引用 85 次
- Model Agnostic Sample Reweighting for Out-of-Distribution LearningXiao Zhou, Yong Lin, Renjie Pi, Weizhong Zhang 等ICML 2022 · 被引用 73 次
- Cross2StrA: Unpaired Cross-lingual Image Captioning with Cross-lingual Cross-modal Structure-pivoted AlignmentShengqiong Wu, Hao Fei, Wei Ji, Tat-Seng ChuaACL 2023 · 被引用 42 次
- Probabilistic Bilevel Coreset SelectionXiao Zhou, Renjie Pi, Weizhong Zhang, Yong Lin 等ICML 2022 · 被引用 39 次
- CgT-GAN: CLIP-guided Text GAN for Image CaptioningJiarui Yu, Haoran Li, Yanbin Hao, Bin Zhu 等ACM MM 2023 · 被引用 26 次
它引用的顶会 Paper4
- Unpaired Image Captioning via Scene Graph AlignmentsJiuxiang Gu, Shafiq R. Joty, Jianfei Cai, Handong Zhao 等ICCV 2019 · 被引用 191 次
- Towards Unsupervised Image Captioning With Shared Multimodal EmbeddingsIro Laina, Christian Rupprecht, Nassir NavabICCV 2019 · 被引用 115 次
- In Defense of Scene Graphs for Image CaptioningKien Nguyen, Subarna Tripathi, Bang Du, Tanaya Guha 等ICCV 2021 · 被引用 55 次
- Self-Supervised Relationship ProbingJiuxiang Gu, Jason Kuen, Shafiq R. Joty, Jianfei Cai 等NeurIPS 2020 · 被引用 18 次
相关 Paper
- Improving Cross-Modal Alignment with Synthetic Pairs for Text-Only Image CaptioningZhiyue Liu, Jinyuan Liu, Fanrong MaAAAI 2024 · 被引用 23 次
- MAGIC: Multimodal relAtional Graph adversarIal inferenCe for Diverse and Unpaired Text-Based Image CaptioningWenqiao Zhang, Haochen Shi, Jiannan Guo, Shengyu Zhang 等AAAI 2022 · 被引用 52 次
- Detach and Attach: Stylized Image Captioning without Paired Stylized DatasetYutong Tan, Zheng Lin, Peng Fu, Mingyu Zheng 等ACM MM 2022 · 被引用 8 次
- Relational Distant Supervision for Image Captioning without Image-Text PairsYayun Qi, Wentian Zhao, Xinxiao WuAAAI 2024 · 被引用 5 次
- Leveraging Unpaired Data for Vision-Language Generative Models via Cycle ConsistencyTianhong Li, Sangnie Bhardwaj, Yonglong Tian, Han Zhang 等ICLR 2024 · 被引用 8 次
