Towards Fully Automated Manga Translation
Ryota Hinami, Shonosuke Ishiwatari, Kazuhiko Yasuda, Yusuke Matsui
摘要
We tackle the problem of machine translation of manga, Japanese comics. Manga translation involves two important problems in machine translation: context-aware and multimodal translation. Since text and images are mixed up in an unstructured fashion in Manga, obtaining context from the image is essential for manga translation. However, it is still an open problem how to extract context from image and integrate into MT models. In addition, corpus and benchmarks to train and evaluate such model is currently unavailable. In this paper, we make the following four contributions that establishes the foundation of manga translation research. First, we propose multimodal context-aware translation framework. We are the first to incorporate context information obtained from manga image. It enables us to translate texts in speech bubbles that cannot be translated without using context information (e.g., texts in other speech bubbles, gender of speakers, etc.). Second, for training the model, we propose the approach to automatic corpus construction from pairs of original manga and their translations, by which large parallel corpus can be constructed without any manual labeling. Third, we created a new benchmark to evaluate manga translation. Finally, on top of our proposed methods, we devised a first compleheisive system for fully automated manga translation.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- The Manga Whisperer: Automatically Generating Transcriptions for ComicsRagav Sachdeva, Andrew ZissermanCVPR 2024 · 被引用 11 次
- Single-to-mix Modality Alignment with Multimodal Large Language Model for Document Image Machine TranslationYupu Liang, Yaping Zhang, Zhiyang Zhang, Yang Zhao 等ACL 2025 · 被引用 6 次
- MMTIT-Bench: A Multilingual and Multi-Scenario Benchmark with Cognition-Perception-Reasoning Guided Text-Image Machine TranslationGengluo Li, Chengquan Zhang, Yupu Liang, Huawen Shen 等CVPR 2026 · 被引用 6 次
- Zero-Shot Character Identification and Speaker Prediction in Comics via Iterative Multimodal FusionYingxuan Li, Ryota Hinami, Kiyoharu Aizawa, Yusuke MatsuiACM MM 2024 · 被引用 3 次
- MT³: A Synergistic Multi-Task RL Framework for Specializing MLLMs in Text Image Machine TranslationZhaopeng Feng, Yupu Liang, Shaosheng Cao, Jiayuan Su 等ACL 2026
它引用的顶会 Paper1
相关 Paper
- Seamless manga inpainting with semantics awarenessMinshan Xie, Menghan Xia, Xueting Liu, Chengze Li 等SIGGRAPH 2021 · 被引用 29 次
- MangaGAN: Unpaired Photo-to-Manga Translation Based on The Methodology of Manga DrawingHao Su, Jianwei Niu, Xuefeng Liu, Qingfeng Li 等AAAI 2021 · 被引用 37 次
- Enhancing Entertainment Translation for Indian Languages Using Adaptive Context, Style and LLMsPratik Rakesh Singh, Mohammadi Zaki, Pankaj WasnikAAAI 2025 · 被引用 2 次
- LVP-M3: Language-aware Visual Prompt for Multilingual Multimodal Machine TranslationHongcheng Guo, Jiaheng Liu, Haoyang Huang, Jian Yang 等EMNLP 2022 · 被引用 9 次
- Region-Wise Correspondence Prediction between Manga Line Art ImagesYingxuan Li, Jiafeng Mao, Qianru Qiu, Yusuke MatsuiCVPR 2026
