Product-oriented Machine Translation with Cross-modal Cross-lingual Pre-training
Yuqing Song, Shizhe Chen, Qin Jin, Wei Luo, Jun Xie, Fei Huang
Abstract
Translating e-commercial product descriptions, a.k.a product-oriented machine translation (PMT), is essential to serve e-shoppers all over the world. However, due to the domain specialty, the PMT task is more challenging than traditional machine translation problems. Firstly, there are many specialized jargons in the product description, which are ambiguous to translate without the product image. Secondly, product descriptions are related to the image in more complicated ways than standard image descriptions, involving various visual aspects such as objects, shapes, colors or even subjective styles. Moreover, existing PMT datasets are small in scale to support the research. In this paper, we first construct a large-scale bilingual product description dataset called Fashion-MMT, which contains over 114k noisy and 40k manually cleaned description translations with multiple product images. To effectively learn semantic alignments among product images and bilingual texts in translation, we design a unified product-oriented cross-modal cross-lingual model (UPOC 2 ) for pre-training and fine-tuning. Experiments on the Fashion-MMT and Multi30k datasets show that our model significantly outperforms the state-of-the-art models even pre-trained on the same dataset. It is also shown to benefit more from large-scale noisy data to improve the translation quality. We will release the dataset and codes at https://github.com/syuqings/Fashion-MMT.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 50deb370-a0fa-4259-a500-a3b45bbd9774Cited by top-tier papers8
- Cross-Lingual Cross-Modal Retrieval with Noise-Robust LearningYabing Wang, Jianfeng Dong, Tianxiang Liang, Minsong Zhang et al.ACM MM 2022 · 26 citations
- Exploring Better Text Image Translation with Multimodal CodebookZhibin Lan, Jiawei Yu, Xiang Li, Wen Zhang et al.ACL 2023 · 12 citations
- PEIT: Bridging the Modality Gap with Pre-trained Models for End-to-End Image TranslationShaolin Zhu, Shangjie Li, Yikun Lei, Deyi XiongACL 2023 · 12 citations
- Increasing Visual Awareness in Multimodal Neural Machine Translation from an Information Theoretic PerspectiveBaijun Ji, Tong Zhang, Yicheng Zou, Bojie Hu et al.EMNLP 2022 · 11 citations
- Cultural Concept Adaptation on Multimodal ReasoningZhi Li, Yin ZhangEMNLP 2023 · 6 citations
Builds on13
- Unified Vision-Language Pre-Training for Image Captioning and VQALuowei Zhou, Hamid Palangi, Lei Zhang, Houdong Hu et al.AAAI 2020 · 1,047 citations
- Unicoder-VL: A Universal Encoder for Vision and Language by Cross-Modal Pre-TrainingGen Li, Nan Duan, Yuejian Fang, Ming Gong et al.AAAI 2020 · 966 citations
- Unsupervised Cross-lingual Representation Learning at ScaleAlexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary et al.ACL 2020 · 539 citations
- A Novel Graph-based Multi-modal Fusion Encoder for Neural Machine TranslationYongjing Yin, Fandong Meng, Jinsong Su, Chulun Zhou et al.ACL 2020 · 145 citations
- Neural Machine Translation with Universal Visual RepresentationZhuosheng Zhang, Kehai Chen, Rui Wang, Masao Utiyama et al.ICLR 2020 · 117 citations
Related papers
- Cross-Lingual Low-Resource Set-to-Description Retrieval for Global E-CommerceJuntao Li, Chang Liu, Jian Wang, Lidong Bing et al.AAAI 2020 · 14 citations
- Visually Precise QueryRiddhiman Dasgupta, Francis Tom, Sudhir Kumar, Mithun Das Gupta et al.ACM MM 2020 · 1 citation
- Seeing the Abstract: Translating the Abstract Language for Vision Language ModelsDavide Talon, Federico Girella, Ziyue Liu, Marco Cristani et al.CVPR 2025
- PRIM: Towards Practical In-Image Multilingual Machine TranslationYanzhi Tian, Zeming Liu, Zhengyang Liu, Chong Feng et al.EMNLP 2025 · 4 citations
- MCM-DPO: Multifaceted Cross-Modal Direct Preference Optimization for Alt-text GenerationJinlan Fu, Shenzhen Huangfu, Hao Fei, Yichong Huang et al.ACM MM 2025
