The Manga Whisperer: Automatically Generating Transcriptions for Comics
Ragav Sachdeva, Andrew Zisserman
Abstract
In the past few decades, Japanese comics, commonly referred to as Manga, have transcended both cultural and linguistic boundaries to become a true worldwide sensation. Yet, the inherent reliance on visual cues and illustration within manga renders it largely inaccessible to individuals with visual impairments. In this work, we seek to address this substantial barrier, with the aim of ensuring that manga can be appreciated and actively engaged by everyone. Specifically, we tackle the problem of diarisation i.e. generating a transcription of who said what and when, in a fully automatic way. To this end, we make the following contributions: (1) we present a unified model, Magi, that is able to (a) detect panels, text boxes and character boxes, (b) cluster characters by identity (without knowing the number of clusters apriori), and (c) associate dialogues to their speakers; (2) we propose a novel approach that is able to sort the detected text boxes in their reading order and generate a dialogue transcript; (3) we annotate an evaluation benchmark for this task using publicly available [English] manga pages. The code, evaluation datasets and the pretrained model can be found at: https://github.com/ragavsachdeva/magi.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers4
- DreamingComics: A Story Visualization Pipeline via Subject and Layout Customized Generation using Video ModelsPatrick Kwon, Chen ChenCVPR 2026 · 1 citation
- From Panels to Prose: Generating Literary Narratives from ComicsRagav Sachdeva, Andrew ZissermanICCV 2025 · 1 citation
- DiffSensei: Bridging Multi-Modal LLMs and Diffusion Models for Customized Manga GenerationJianzong Wu, Chao Tang, Jingbo Wang, Yanhong Zeng et al.CVPR 2025
- Advancing Manga Analysis: Comprehensive Segmentation Annotations for the Manga109 DatasetMinshan Xie, Jian Lin, Hanyuan Liu, Chengze Li et al.CVPR 2025
Builds on9
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Deformable DETR: Deformable Transformers for End-to-End Object DetectionXizhou Zhu, Weijie Su, Lewei Lu, Bin Li et al.ICLR 2021 · 7,353 citations
- Supervised Contrastive LearningPrannay Khosla, Piotr Teterwak, Chen Wang, Aaron Sarna et al.NeurIPS 2020 · 7,049 citations
Related papers
- Towards Fully Automated Manga TranslationRyota Hinami, Shonosuke Ishiwatari, Kazuhiko Yasuda, Yusuke MatsuiAAAI 2021 · 41 citations
- Seamless manga inpainting with semantics awarenessMinshan Xie, Menghan Xia, Xueting Liu, Chengze Li et al.SIGGRAPH 2021 · 29 citations
- AutoAD II: The Sequel - Who, When, and What in Movie Audio DescriptionTengda Han, Max Bain, Arsha Nagrani, Gül Varol et al.ICCV 2023 · 55 citations
- Multimodal Persona Based Generation of Comic DialogsHarsh Agrawal, Aditya Mishra, Manish Gupta, MausamACL 2023 · 8 citations
- CineSRD: Leveraging Visual, Acoustic, and Linguistic Cues for Open-World Visual Media Speaker DiarizationLiangbin Huang, Xiaohua Liao, Chaoqun Cui, Shijing Wang et al.CVPR 2026
