Exploring Better Text Image Translation with Multimodal Codebook
Zhibin Lan, Jiawei Yu, Xiang Li, Wen Zhang, Jian Luan, Bin Wang, Degen Huang, Jinsong Su
Abstract
Text image translation (TIT) aims to translate the source texts embedded in the image to target translations, which has a wide range of applications and thus has important research value. However, current studies on TIT are confronted with two main bottlenecks: 1) this task lacks a publicly available TIT dataset, 2) dominant models are constructed in a cascaded manner, which tends to suffer from the error propagation of optical character recognition (OCR). In this work, we first annotate a Chinese-English TIT dataset named OCRMT30K, providing convenience for subsequent studies. Then, we propose a TIT model with a multimodal codebook, which is able to associate the image with relevant texts, providing useful supplementary information for translation. Moreover, we present a multi-stage training framework involving text machine translation, image-text alignment, and TIT tasks, which fully exploits additional bilingual texts, OCR dataset and our OCRMT30K dataset to train our model. Extensive experiments and in-depth analyses strongly demonstrate the effectiveness of our proposed model and training framework. 1
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 672bc005-dd9e-410c-bf23-3926f57a9feeCited by top-tier papers7
- Single-to-mix Modality Alignment with Multimodal Large Language Model for Document Image Machine TranslationYupu Liang, Yaping Zhang, Zhiyang Zhang, Yang Zhao et al.ACL 2025 · 6 citations
- MMTIT-Bench: A Multilingual and Multi-Scenario Benchmark with Cognition-Perception-Reasoning Guided Text-Image Machine TranslationGengluo Li, Chengquan Zhang, Yupu Liang, Huawen Shen et al.CVPR 2026 · 6 citations
- PRIM: Towards Practical In-Image Multilingual Machine TranslationYanzhi Tian, Zeming Liu, Zhengyang Liu, Chong Feng et al.EMNLP 2025 · 4 citations
- Towards Better Multi-modal Keyphrase Generation via Visual Entity Enhancement and Multi-granularity Image Noise FilteringYifan Dong, Suhang Wu, Fandong Meng, Jie Zhou et al.ACM MM 2023 · 3 citations
- VaaWIT: Visual-Aware Adaptation of Large Language Models for Multilingual Web Image TranslationBo Li, Ronghao Chen, Ningyuan Deng, Huacan Wang et al.KDD 2026
Builds on12
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- VaTeX: A Large-Scale, High-Quality Multilingual Dataset for Video-and-Language ResearchXin Wang, Jiawei Wu, Jun-Kun Chen, Lei Li et al.ICCV 2019 · 688 citations
- A Novel Graph-based Multi-modal Fusion Encoder for Neural Machine TranslationYongjing Yin, Fandong Meng, Jinsong Su, Chulun Zhou et al.ACL 2020 · 145 citations
- Neural Machine Translation with Universal Visual RepresentationZhuosheng Zhang, Kehai Chen, Rui Wang, Masao Utiyama et al.ICLR 2020 · 117 citations
- On Vision Features in Multimodal Machine TranslationBei Li, Chuanhao Lv, Zefan Zhou, Tao Zhou et al.ACL 2022 · 82 citations
Related papers
- PEIT: Bridging the Modality Gap with Pre-trained Models for End-to-End Image TranslationShaolin Zhu, Shangjie Li, Yikun Lei, Deyi XiongACL 2023 · 12 citations
- MT³: A Synergistic Multi-Task RL Framework for Specializing MLLMs in Text Image Machine TranslationZhaopeng Feng, Yupu Liang, Shaosheng Cao, Jiayuan Su et al.ACL 2026
- Translation-Enhanced Multilingual Text-to-Image GenerationYaoyiran Li, Ching-Yun Chang, Stephen Rawls, Ivan Vulic et al.ACL 2023 · 8 citations
- Zero-TextCap: Zero-shot Framework for Text-based Image CaptioningDongsheng Xu, Wenye Zhao, Yi Cai, Qingbao HuangACM MM 2023 · 4 citations
- TAP: Text-Aware Pre-Training for Text-VQA and Text-CaptionZhengyuan Yang, Yijuan Lu, Jianfeng Wang, Xi Yin et al.CVPR 2021
