Cross2StrA: Unpaired Cross-lingual Image Captioning with Cross-lingual Cross-modal Structure-pivoted Alignment
Shengqiong Wu, Hao Fei, Wei Ji, Tat-Seng Chua
Abstract
Unpaired cross-lingual image captioning has long suffered from irrelevancy and disfluency issues, due to the inconsistencies of the semantic scene and syntax attributes during transfer. In this work, we propose to address the above problems by incorporating the scene graph (SG) structures and the syntactic constituency (SC) trees. Our captioner contains the semantic structure-guided image-to-pivot captioning and the syntactic structure-guided pivot-to-target translation, two of which are joined via pivot language. We then take the SG and SC structures as pivoting, performing cross-modal semantic structure alignment and cross-lingual syntactic structure alignment learning. We further introduce cross-lingual&cross-modal back-translation training to fully align the captioning and translation stages. Experiments on English-Chinese transfers show that our model shows great superiority in improving captioning relevancy and fluency.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext cdc38ba3-c5a7-4aca-94fc-48c6dfa5b5c7Cited by top-tier papers7
- Video-of-Thought: Step-by-Step Video Reasoning from Perception to CognitionHao Fei, Shengqiong Wu, Wei Ji, Hanwang Zhang et al.ICML 2024 · 182 citations
- Semi-Supervised Panoptic Narrative GroundingDanni Yang, Jiayi Ji, Xiaoshuai Sun, Haowei Wang et al.ACM MM 2023 · 7 citations
- Improving Panoptic Narrative Grounding by Harnessing Semantic Relationships and Visual ConfirmationTianyu Guo, Haowei Wang, Yiwei Ma, Jiayi Ji et al.AAAI 2024 · 5 citations
- Synergistic Dual Spatial-aware Generation of Image-to-text and Text-to-imageYu Zhao, Hao Fei, Xiangtai Li, Libo Qin et al.NeurIPS 2024 · 2 citations
- SpeechEE: A Novel Benchmark for Speech Event ExtractionBin Wang, Meishan Zhang, Hao Fei, Yu Zhao et al.ACM MM 2024 · 1 citation
Builds on10
- Improving Image Captioning by Leveraging Intra- and Inter-layer Global Representation in Transformer NetworkJiayi Ji, Yunpeng Luo, Xiaoshuai Sun, Fuhai Chen et al.AAAI 2021 · 206 citations
- LasUIE: Unifying Information Extraction with Latent Adaptive Structure-aware Generative Language ModelHao Fei, Shengqiong Wu, Jingye Li, Bobo Li et al.NeurIPS 2022 · 114 citations
- Encoder-Decoder Based Unified Semantic Role Labeling with Label-Aware SyntaxHao Fei, Fei Li, Bobo Li, Donghong JiAAAI 2021 · 60 citations
- Unsupervised Multimodal Neural Machine Translation with Pseudo Visual PivotingPo-Yao Huang, Junjie Hu, Xiaojun Chang, Alexander G. HauptmannACL 2020 · 43 citations
- Learning Source Phrase Representations for Neural Machine TranslationHongfei Xu, Josef van Genabith, Deyi Xiong, Qiuhui Liu et al.ACL 2020 · 18 citations
Related papers
- Unpaired Image Captioning via Scene Graph AlignmentsJiuxiang Gu, Shafiq R. Joty, Jianfei Cai, Handong Zhao et al.ICCV 2019 · 191 citations
- UNISON: Unpaired Cross-Lingual Image CaptioningJiahui Gao, Yi Zhou, Philip L. H. Yu, Shafiq R. Joty et al.AAAI 2022 · 18 citations
- FAIEr: Fidelity and Adequacy Ensured Image Caption EvaluationSijin Wang, Ziwei Yao, Ruiping Wang, Zhongqin Wu et al.CVPR 2021
- Triangle-Reward Reinforcement Learning: A Visual-Linguistic Semantic Alignment for Image CaptioningWeizhi Nie, Jiesi Li, Ning Xu, An-An Liu et al.ACM MM 2021 · 9 citations
- In Defense of Scene Graphs for Image CaptioningKien Nguyen, Subarna Tripathi, Bang Du, Tanaya Guha et al.ICCV 2021 · 55 citations
