DiffSensei: Bridging Multi-Modal LLMs and Diffusion Models for Customized Manga Generation
Jianzong Wu, Chao Tang, Jingbo Wang, Yanhong Zeng, Xiangtai Li, Yunhai Tong
摘要
Story visualization, the task of creating visual narratives from textual descriptions, has seen progress with text-toimage generation models. However, these models often lack effective control over character appearances and interactions, particularly in multi-character scenes. To address these limitations, we propose a new task: customized manga generation and introduce DiffSensei, an innovative framework specifically designed for generating manga with dynamic multi-character control. DiffSensei integrates a diffusion-based image generator with a multimodal large language model (MLLM) that acts as a text-compatible identity adapter. Our approach employs masked crossattention to seamlessly incorporate character features, enabling precise layout control without direct pixel transfer. Additionally, the MLLM-based adapter adjusts character features to align with panel-specific text cues, allowing flex-ible adjustments in character expressions, poses, and actions. We also introduce MangaZero, a large-scale dataset tailored to this task, containing 43,264 manga pages and 427,147 annotated panels, supporting the visualization of varied character interactions and movements across sequential frames. Extensive experiments demonstrate that DiffSensei outperforms existing models, marking a significant advancement in manga generation by enabling textadaptable character customization. The code, model, and dataset are open-sourced to the community. 1
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- ViStoryBench: Comprehensive Benchmark Suite for Story VisualizationCailin Zhuang, Ailin Huang, Hu Yaoqi, Jingwei Wu 等CVPR 2026 · 被引用 37 次
- Story-Iter: A Training-free Iterative Paradigm for Long Story VisualizationJiawei Mao, Xiaoke Huang, Yunfei Xie, Yuanqi Chang 等ICLR 2026 · 被引用 18 次
- LogiStory: A Logic-Aware Framework for Multi-Image Story VisualizationChutian Meng, Fan Ma, Chi Zhang, Jiaxu Miao 等ICLR 2026 · 被引用 3 次
- DreamingComics: A Story Visualization Pipeline via Subject and Layout Customized Generation using Video ModelsPatrick Kwon, Chen ChenCVPR 2026 · 被引用 1 次
- StoryTailor: A Zero-Shot Pipeline for Action-Rich Multi-Subject Visual NarrativesJinghao Hu, Yuhe Zhang, Guohua Geng, Kang Li 等CVPR 2026
它引用的顶会 Paper25
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 被引用 11,349 次
- Adding Conditional Control to Text-to-Image Diffusion ModelsLvmin Zhang, Anyi Rao, Maneesh AgrawalaICCV 2023 · 被引用 6,759 次
- SDXL: Improving Latent Diffusion Models for High-Resolution Image SynthesisDustin Podell, Zion English, Kyle Lacey, Andreas Blattmann 等ICLR 2024 · 被引用 4,569 次
相关 Paper
- Storynizor: Consistent Story Generation via Inter-Frame Synchronized and Shuffled ID InjectionYuhang Ma, Wenting Xu, Chaoyi Zhao, Keqiang Sun 等AAAI 2025 · 被引用 1 次
- From Panels to Prose: Generating Literary Narratives from ComicsRagav Sachdeva, Andrew ZissermanICCV 2025 · 被引用 1 次
- MVCustom: Multi-View Customized Diffusion via Geometric Latent Rendering and CompletionMinjung Shin, Hyunin Cho, Sooyeon Go, Jin-Hwa Kim 等ICLR 2026 · 被引用 3 次
- MultiDiffusion: Fusing Diffusion Paths for Controlled Image GenerationOmer Bar-Tal, Lior Yariv, Yaron Lipman, Tali DekelICML 2023 · 被引用 575 次
- The Manga Whisperer: Automatically Generating Transcriptions for ComicsRagav Sachdeva, Andrew ZissermanCVPR 2024 · 被引用 11 次
