DiffSensei: Bridging Multi-Modal LLMs and Diffusion Models for Customized Manga Generation
Jianzong Wu, Chao Tang, Jingbo Wang, Yanhong Zeng, Xiangtai Li, Yunhai Tong
Abstract
Story visualization, the task of creating visual narratives from textual descriptions, has seen progress with text-toimage generation models. However, these models often lack effective control over character appearances and interactions, particularly in multi-character scenes. To address these limitations, we propose a new task: customized manga generation and introduce DiffSensei, an innovative framework specifically designed for generating manga with dynamic multi-character control. DiffSensei integrates a diffusion-based image generator with a multimodal large language model (MLLM) that acts as a text-compatible identity adapter. Our approach employs masked crossattention to seamlessly incorporate character features, enabling precise layout control without direct pixel transfer. Additionally, the MLLM-based adapter adjusts character features to align with panel-specific text cues, allowing flex-ible adjustments in character expressions, poses, and actions. We also introduce MangaZero, a large-scale dataset tailored to this task, containing 43,264 manga pages and 427,147 annotated panels, supporting the visualization of varied character interactions and movements across sequential frames. Extensive experiments demonstrate that DiffSensei outperforms existing models, marking a significant advancement in manga generation by enabling textadaptable character customization. The code, model, and dataset are open-sourced to the community. 1
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers5
- ViStoryBench: Comprehensive Benchmark Suite for Story VisualizationCailin Zhuang, Ailin Huang, Hu Yaoqi, Jingwei Wu et al.CVPR 2026 · 37 citations
- Story-Iter: A Training-free Iterative Paradigm for Long Story VisualizationJiawei Mao, Xiaoke Huang, Yunfei Xie, Yuanqi Chang et al.ICLR 2026 · 18 citations
- LogiStory: A Logic-Aware Framework for Multi-Image Story VisualizationChutian Meng, Fan Ma, Chi Zhang, Jiaxu Miao et al.ICLR 2026 · 3 citations
- DreamingComics: A Story Visualization Pipeline via Subject and Layout Customized Generation using Video ModelsPatrick Kwon, Chen ChenCVPR 2026 · 1 citation
- StoryTailor: A Zero-Shot Pipeline for Action-Rich Multi-Subject Visual NarrativesJinghao Hu, Yuhe Zhang, Guohua Geng, Kang Li et al.CVPR 2026
Builds on25
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
- Adding Conditional Control to Text-to-Image Diffusion ModelsLvmin Zhang, Anyi Rao, Maneesh AgrawalaICCV 2023 · 6,759 citations
- SDXL: Improving Latent Diffusion Models for High-Resolution Image SynthesisDustin Podell, Zion English, Kyle Lacey, Andreas Blattmann et al.ICLR 2024 · 4,569 citations
Related papers
- Storynizor: Consistent Story Generation via Inter-Frame Synchronized and Shuffled ID InjectionYuhang Ma, Wenting Xu, Chaoyi Zhao, Keqiang Sun et al.AAAI 2025 · 1 citation
- From Panels to Prose: Generating Literary Narratives from ComicsRagav Sachdeva, Andrew ZissermanICCV 2025 · 1 citation
- MVCustom: Multi-View Customized Diffusion via Geometric Latent Rendering and CompletionMinjung Shin, Hyunin Cho, Sooyeon Go, Jin-Hwa Kim et al.ICLR 2026 · 3 citations
- MultiDiffusion: Fusing Diffusion Paths for Controlled Image GenerationOmer Bar-Tal, Lior Yariv, Yaron Lipman, Tali DekelICML 2023 · 575 citations
- The Manga Whisperer: Automatically Generating Transcriptions for ComicsRagav Sachdeva, Andrew ZissermanCVPR 2024 · 11 citations
