From Panels to Prose: Generating Literary Narratives from Comics
Ragav Sachdeva, Andrew Zisserman
Abstract
Comics have long been a popular form of storytelling, offering visually engaging narratives that captivate audiences worldwide. However, the visual nature of comics presents a significant barrier for visually impaired readers, limiting their access to these engaging stories. In this work, we provide a pragmatic solution to this accessibility challenge by developing an automated system that generates text-based literary narratives from manga comics. Our approach aims to create an evocative and immersive prose that not only conveys the original narrative but also captures the depth and complexity of characters, their interactions, and the vivid settings in which they reside. To this end we make the following contributions: (1) We present a unified model, Magiv3, that excels at various functional tasks pertaining to comic understanding, such as localising panels, characters, texts, and speech-bubble tails, performing OCR, grounding characters etc. (2) We release human-annotated captions for over 3300 Japanese comic panels, along with character grounding annotations, and benchmark large vision-language models in their ability to understand comic images. (3) Finally, we demonstrate how integrating large vision-language models with Magiv3, can generate seamless literary narratives that allows visually impaired audiences to engage with the depth and richness of comic storytelling.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on9
- InstructBLIP: Towards General-purpose Vision-Language Models with Instruction TuningWenliang Dai, Junnan Li, Dongxu Li, Anthony Meng Huat Tiong et al.NeurIPS 2023 · 4,013 citations
- Pix2seq: A Language Modeling Framework for Object DetectionTing Chen, Saurabh Saxena, Lala Li, David J. Fleet et al.ICLR 2022 · 435 citations
- A Unified Sequence Interface for Vision TasksTing Chen, Saurabh Saxena, Lala Li, Tsung-Yi Lin et al.NeurIPS 2022 · 201 citations
- Decoupled Adaptation for Cross-Domain Object DetectionJunguang Jiang, Baixu Chen, Jianmin Wang, Mingsheng LongICLR 2022 · 88 citations
- The Manga Whisperer: Automatically Generating Transcriptions for ComicsRagav Sachdeva, Andrew ZissermanCVPR 2024 · 11 citations
Related papers
- DiffSensei: Bridging Multi-Modal LLMs and Diffusion Models for Customized Manga GenerationJianzong Wu, Chao Tang, Jingbo Wang, Yanhong Zeng et al.CVPR 2025
- Movie101v2: Improved Movie Narration BenchmarkZihao Yue, Yepeng Zhang, Ziheng Wang, Qin JinACL 2025 · 5 citations
- GenAssist: Making Image Generation AccessibleMina Huh, Yi-Hao Peng, Amy PavelUIST 2023 · 58 citations
- Towards Fully Automated Manga TranslationRyota Hinami, Shonosuke Ishiwatari, Kazuhiko Yasuda, Yusuke MatsuiAAAI 2021 · 41 citations
- VisText: A Benchmark for Semantically Rich Chart CaptioningBenny J. Tang, Angie W. Boggust, Arvind SatyanarayanACL 2023 · 44 citations
