MSCTD: A Multimodal Sentiment Chat Translation Dataset
Yunlong Liang, Fandong Meng, Jinan Xu, Yufeng Chen, Jie Zhou
Abstract
Multimodal machine translation and textual chat translation have received considerable attention in recent years. Although the conversation in its natural form is usually multimodal, there still lacks work on multimodal machine translation in conversations. In this work, we introduce a new task named Multimodal Chat Translation (MCT), aiming to generate more accurate translations with the help of the associated dialogue history and visual context. To this end, we firstly construct a Multimodal Sentiment Chat Translation Dataset (MSCTD) containing 142,871 English-Chinese utterance pairs in 14,762 bilingual dialogues and 30,370 English-German utterance pairs in 3,079 bilingual dialogues. Each utterance pair, corresponding to the visual context that reflects the current conversational scene, is annotated with a sentiment label. Then, we benchmark the task by establishing multiple baseline systems that incorporate multimodal and sentiment features for MCT. Preliminary experiments on four language directions (English↔Chinese and English↔German) verify the potential of contextual and multimodal information fusion and the positive impact of sentiment on the MCT task. Additionally, as a by-product of the MSCTD, it also provides two new benchmarks on multimodal dialogue sentiment analysis. Our work can facilitate research on both multimodal chat translation and multimodal dialogue sentiment analysis. 1
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 51c0145e-8d14-4818-ae29-72ac28b6b0bfCited by top-tier papers7
- AfriSenti: A Twitter Sentiment Analysis Benchmark for African LanguagesShamsuddeen Hassan Muhammad, Idris Abdulmumin, Abinew Ali Ayele, Nedjma Ousidhoum et al.EMNLP 2023 · 33 citations
- Summary-Oriented Vision Modeling for Multimodal Abstractive SummarizationYunlong Liang, Fandong Meng, Jinan Xu, Jiaan Wang et al.ACL 2023 · 17 citations
- Scheduled Multi-task Learning for Neural Chat TranslationYunlong Liang, Fandong Meng, Jinan Xu, Yufeng Chen et al.ACL 2022 · 15 citations
- Bridging the Gap between Synthetic and Authentic Images for Multimodal Machine TranslationWenyu Guo, Qingkai Fang, Dong Yu, Yang FengEMNLP 2023 · 5 citations
- CrisisTS: Coupling Social Media Textual Data and Meteorological Time Series for Urgency ClassificationRomain Meunier, Farah Benamara, Véronique Moriceau, Zhongzheng Qiao et al.ACL 2025 · 1 citation
Builds on6
- Unsupervised Cross-lingual Representation Learning at ScaleAlexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary et al.ACL 2020 · 539 citations
- A Novel Graph-based Multi-modal Fusion Encoder for Neural Machine TranslationYongjing Yin, Fandong Meng, Jinsong Su, Chulun Zhou et al.ACL 2020 · 145 citations
- Infusing Multi-Source Knowledge with Heterogeneous Graph Neural Network for Emotional Conversation GenerationYunlong Liang, Fandong Meng, Ying Zhang, Yufeng Chen et al.AAAI 2021 · 62 citations
- Dynamic Context-guided Capsule Network for Multimodal Machine TranslationHuan Lin, Fandong Meng, Jinsong Su, Yongjing Yin et al.ACM MM 2020 · 57 citations
- Towards Making the Most of Dialogue Characteristics for Neural Chat TranslationYunlong Liang, Chulun Zhou, Fandong Meng, Jinan Xu et al.EMNLP 2021 · 12 citations
Related papers
- LVP-M3: Language-aware Visual Prompt for Multilingual Multimodal Machine TranslationHongcheng Guo, Jiaheng Liu, Haoyang Huang, Jian Yang et al.EMNLP 2022 · 9 citations
- Let's Go Real Talk: Spoken Dialogue Model for Face-to-Face ConversationSe Jin Park, Chae Won Kim, Hyeongseop Rha, Minsu Kim et al.ACL 2024
- Impact of Stickers on Multimodal Sentiment and Intent in Social Media: A New Task, Dataset and BaselineYuanchen Shi, Fang Kong, Longyin ZhangACM MM 2025 · 3 citations
- Multimodal Multi-turn Conversation Stance Detection: A Challenge Dataset and Effective ModelFuqiang Niu, Zebang Cheng, Xianghua Fu, Xiaojiang Peng et al.ACM MM 2024 · 13 citations
- Modeling Bilingual Conversational Characteristics for Neural Chat TranslationYunlong Liang, Fandong Meng, Yufeng Chen, Jinan Xu et al.ACL 2021
