Soul-Mix: Enhancing Multimodal Machine Translation with Manifold Mixup
Xuxin Cheng, Ziyu Yao, Yifei Xin, Hao An, Hongxiang Li, Yaowei Li, Yuexian Zou
摘要
Multimodal machine translation (MMT) aims to improve the performance of machine translation with the help of visual information, which has received widespread attention recently. It has been verified that visual information brings greater performance gains when the textual information is limited. However, most previous works ignore to take advantage of the complete textual inputs and the limited textual inputs at the same time, which limits the overall performance. To solve this issue, we propose a mixup method termed Soul-Mix to enhance MMT by using visual information more effectively. We mix the predicted translations of complete textual input and the limited textual inputs. Experimental results on the Multi30K dataset of three translation directions show that our Soul-Mix significantly outperforms existing approaches and achieves new state-of-the-art performance with fewer parameters than some previous models. Besides, the strength of Soul-Mix is more obvious on more challenging MSCOCO dataset which includes more out-of-domain instances with lots of ambiguous verbs.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- Multimodal Neural Machine Translation: A Survey of the State of the ArtYi Feng, Chuanyi Li, Jiatong He, Zhenyu Hou 等EMNLP 2025 · 被引用 1 次
- SHIFT: Selected Helpful Informative Frame for Video-guided Machine TranslationBoyu Guan, Chuang Han, Yining Zhang, Yupu Liang 等EMNLP 2025
- Scalable Multilingual Multimodal Machine Translation with Speech-Text FusionYexing Du, Youcheng Pan, Zekun Wang, Zheng Chu 等ICLR 2026
- Make Imagination Clearer! Stable Diffusion-based Visual Imagination for Multimodal Machine TranslationAndong Chen, Yuchen Song, Kehai Chen, Xuefeng Bai 等ACL 2025
- DART: Disambiguation-Aware Reasoning for Video-guided Machine TranslationBoyu Guan, Chuang Han, Yang Zhao, Chengqing ZongACL 2026
它引用的顶会 Paper15
- MixText: Linguistically-Informed Interpolation of Hidden Space for Semi-Supervised Text ClassificationJiaao Chen, Zichao Yang, Diyi YangACL 2020 · 被引用 340 次
- Neural Machine Translation with Universal Visual RepresentationZhuosheng Zhang, Kehai Chen, Rui Wang, Masao Utiyama 等ICLR 2020 · 被引用 117 次
- On Vision Features in Multimodal Machine TranslationBei Li, Chuanhao Lv, Zefan Zhou, Tao Zhou 等ACL 2022 · 被引用 82 次
- Enhancing Cross-lingual Transfer by Manifold MixupHuiyun Yang, Huadong Chen, Hao Zhou, Lei LiICLR 2022 · 被引用 49 次
- Scene Graph as Pivoting: Inference-time Image-free Unsupervised Multimodal Machine Translation with Visual Scene HallucinationHao Fei, Qian Liu, Meishan Zhang, Min Zhang 等ACL 2023 · 被引用 45 次
相关 Paper
- Increasing Visual Awareness in Multimodal Neural Machine Translation from an Information Theoretic PerspectiveBaijun Ji, Tong Zhang, Yicheng Zou, Bojie Hu 等EMNLP 2022 · 被引用 11 次
- Visual Agreement Regularized Training for Multi-Modal Machine TranslationPengcheng Yang, Boxing Chen, Pei Zhang, Xu SunAAAI 2020 · 被引用 34 次
- Neural Machine Translation with Phrase-Level Universal Visual RepresentationsQingkai Fang, Yang FengACL 2022
- Unsupervised Multimodal Neural Machine Translation with Pseudo Visual PivotingPo-Yao Huang, Junjie Hu, Xiaojun Chang, Alexander G. HauptmannACL 2020 · 被引用 43 次
- LVP-M3: Language-aware Visual Prompt for Multilingual Multimodal Machine TranslationHongcheng Guo, Jiaheng Liu, Haoyang Huang, Jian Yang 等EMNLP 2022 · 被引用 9 次
