The Manga Whisperer: Automatically Generating Transcriptions for Comics
Ragav Sachdeva, Andrew Zisserman
摘要
In the past few decades, Japanese comics, commonly referred to as Manga, have transcended both cultural and linguistic boundaries to become a true worldwide sensation. Yet, the inherent reliance on visual cues and illustration within manga renders it largely inaccessible to individuals with visual impairments. In this work, we seek to address this substantial barrier, with the aim of ensuring that manga can be appreciated and actively engaged by everyone. Specifically, we tackle the problem of diarisation i.e. generating a transcription of who said what and when, in a fully automatic way. To this end, we make the following contributions: (1) we present a unified model, Magi, that is able to (a) detect panels, text boxes and character boxes, (b) cluster characters by identity (without knowing the number of clusters apriori), and (c) associate dialogues to their speakers; (2) we propose a novel approach that is able to sort the detected text boxes in their reading order and generate a dialogue transcript; (3) we annotate an evaluation benchmark for this task using publicly available [English] manga pages. The code, evaluation datasets and the pretrained model can be found at: https://github.com/ragavsachdeva/magi.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- DreamingComics: A Story Visualization Pipeline via Subject and Layout Customized Generation using Video ModelsPatrick Kwon, Chen ChenCVPR 2026 · 被引用 1 次
- From Panels to Prose: Generating Literary Narratives from ComicsRagav Sachdeva, Andrew ZissermanICCV 2025 · 被引用 1 次
- DiffSensei: Bridging Multi-Modal LLMs and Diffusion Models for Customized Manga GenerationJianzong Wu, Chao Tang, Jingbo Wang, Yanhong Zeng 等CVPR 2025
- Advancing Manga Analysis: Comprehensive Segmentation Annotations for the Manga109 DatasetMinshan Xie, Jian Lin, Hanyuan Liu, Chengze Li 等CVPR 2025
它引用的顶会 Paper9
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 被引用 24,064 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Deformable DETR: Deformable Transformers for End-to-End Object DetectionXizhou Zhu, Weijie Su, Lewei Lu, Bin Li 等ICLR 2021 · 被引用 7,353 次
- Supervised Contrastive LearningPrannay Khosla, Piotr Teterwak, Chen Wang, Aaron Sarna 等NeurIPS 2020 · 被引用 7,049 次
相关 Paper
- Towards Fully Automated Manga TranslationRyota Hinami, Shonosuke Ishiwatari, Kazuhiko Yasuda, Yusuke MatsuiAAAI 2021 · 被引用 41 次
- Seamless manga inpainting with semantics awarenessMinshan Xie, Menghan Xia, Xueting Liu, Chengze Li 等SIGGRAPH 2021 · 被引用 29 次
- AutoAD II: The Sequel - Who, When, and What in Movie Audio DescriptionTengda Han, Max Bain, Arsha Nagrani, Gül Varol 等ICCV 2023 · 被引用 55 次
- Multimodal Persona Based Generation of Comic DialogsHarsh Agrawal, Aditya Mishra, Manish Gupta, MausamACL 2023 · 被引用 8 次
- CineSRD: Leveraging Visual, Acoustic, and Linguistic Cues for Open-World Visual Media Speaker DiarizationLiangbin Huang, Xiaohua Liao, Chaoqun Cui, Shijing Wang 等CVPR 2026
