Towards Cross-Cultural Machine Translation with Retrieval-Augmented Generation from Multilingual Knowledge Graphs
Simone Conia, Daniel Lee, Min Li, Umar Farooq Minhas, Saloni Potdar, Yunyao Li
Abstract
Translating text that contains entity names is a challenging task, as cultural-related references can vary significantly across languages. These variations may also be caused by transcreation, an adaptation process that entails more than transliteration and word-for-word translation. In this paper, we address the problem of cross-cultural translation on two fronts: (i) we introduce XC-Translate, the first large-scale, manually-created benchmark for machine translation that focuses on text that contains potentially culturally-nuanced entity names, and (ii) we propose KG-MT, a novel end-to-end method to integrate information from a multilingual knowledge graph into a neural machine translation model by leveraging a dense retrieval mechanism. Our experiments and analyses show that current machine translation systems and large language models still struggle to translate texts containing entity names, whereas KG-MT outperforms state-of-the-art approaches by a large margin, obtaining a 129% and 62% relative improvement compared to NLLB-200 and GPT-4, respectively.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 5229725a-4bb5-4eb8-afe7-4ea26e251b73Cited by top-tier papers10
- RAVENEA: A Benchmark for Multimodal Retrieval-Augmented Visual Culture UnderstandingJiaang Li, Yifei Yuan, Wenyan Li, Mohammad Aliannejadi et al.ICLR 2026 · 9 citations
- Exploring the Translation Mechanism of Large Language ModelsHongbin Zhang, Kehai Chen, Xuefeng Bai, Xiucheng Li et al.NeurIPS 2025 · 4 citations
- Locate-and-Focus: Enhancing Terminology Translation in Speech Language ModelsSuhang Wu, Jialong Tang, Chengyi Yang, Pei Zhang et al.ACL 2025 · 4 citations
- Incentivizing Parametric Knowledge via Reinforcement Learning with Verifiable Rewards for Cross-Cultural Entity TranslationJiang Zhou, Xiaohu Zhao, Xinwei Wu, Tianyu Dong et al.ACL 2026 · 2 citations
- M4-RAG: A Massive-Scale Multilingual Multi-Cultural Multimodal RAGDavid Anugraha, Patrick Amadeus Irawan, Anshul Singh, En-Shiun Annie Lee et al.CVPR 2026 · 2 citations
Builds on10
- BERTScore: Evaluating Text Generation with BERTTianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger et al.ICLR 2020 · 8,443 citations
- Unsupervised Cross-lingual Representation Learning at ScaleAlexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary et al.ACL 2020 · 539 citations
- Document-Level Machine Translation with Large Language ModelsLongyue Wang, Chenyang Lyu, Tianbo Ji, Zhirui Zhang et al.EMNLP 2023 · 129 citations
- COMET: A Neural Framework for MT EvaluationRicardo Rei, Craig Stewart, Ana C. Farinha, Alon LavieEMNLP 2020 · 6 citations
- Increasing Coverage and Precision of Textual Information in Multilingual Knowledge GraphsSimone Conia, Min Li, Daniel Lee, Umar Farooq Minhas et al.EMNLP 2023 · 3 citations
Related papers
- Culture-Aware Machine Translation in Large Language Models: Benchmarking and InvestigationZekun Yuan, Yangfan Ye, Xiaocheng Feng, Baohang Li et al.ACL 2026 · 2 citations
- LiveCLKTBench: Towards Reliable Evaluation of Cross-Lingual Knowledge Transfer in Multilingual LLMsPei-Fu Guo, Yun-Da Tsai, Chun-Chia Hsu, Kai-Xin Chen et al.ACL 2026 · 1 citation
- XLM-K: Improving Cross-Lingual Language Model Pre-training with Multilingual KnowledgeXiaoze Jiang, Yaobo Liang, Weizhu Chen, Nan DuanAAAI 2022 · 31 citations
- DEEP: DEnoising Entity Pre-training for Neural Machine TranslationJunjie Hu, Hiroaki Hayashi, Kyunghyun Cho, Graham NeubigACL 2022
- Revisiting Commonsense Reasoning in Machine Translation: Training, Evaluation and ChallengeXuebo Liu, Yutong Wang, Derek F. Wong, Runzhe Zhan et al.ACL 2023 · 3 citations
