Storybooth: Training-Free Multi-Subject Consistency for Improved Visual Storytelling
Jaskirat Singh, Junshen K. Chen, Jonas Kohler, Michael F. Cohen
摘要
Training-free consistent text-to-image generation depicting the same subjects across different images is a topic of widespread recent interest. Existing works in this direction predominantly rely on cross-frame self-attention; which improves subjectconsistency by allowing tokens in each frame to pay attention to tokens in other frames during self-attention computation. While useful for single subjects, we find that it struggles when scaling to multiple characters. In this work, we first analyze the reason for these limitations. Our exploration reveals that the primaryissue stems from self-attention leakage, which is exacerbated when trying to ensure consistency across multiple-characters. This happens when tokens from one subject pay attention to other characters, causing them to appear like each other (e.g., a dog appearing like a duck). Motivated by these findings, we propose StoryBooth: a training-free approach for improving multi-character consistency. In particular, we first leverage multi-modal chain-of-thought reasoning and region-based generation to apriori localize the different subjects across the desired story outputs. The final outputs are then generated using a modified diffusion model which consists of two novel layers: 1) a bounded cross-frame self-attention layer for reducing intercharacter attention leakage, and 2) token-merging layer for improving consistency of fine-grain subject details. Through both qualitative and quantitative results we find that the proposed approach surpasses prior state-of-the-art, exhibiting improved consistency across both multiple-characters and fine-grain subject details.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Lay2Story: Extending Diffusion Transformers for Layout-Togglable Story GenerationAo Ma, Jiasong Feng, Ke Cao, Jing Wang 等ICCV 2025 · 被引用 13 次
- LogiStory: A Logic-Aware Framework for Multi-Image Story VisualizationChutian Meng, Fan Ma, Chi Zhang, Jiaxu Miao 等ICLR 2026 · 被引用 3 次
它引用的顶会 Paper19
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Scalable Diffusion Models with TransformersWilliam Peebles, Saining XieICCV 2023 · 被引用 5,568 次
- AnimateDiff: Animate Your Personalized Text-to-Image Diffusion Models without Specific TuningYuwei Guo, Ceyuan Yang, Anyi Rao, Zhengyang Liang 等ICLR 2024 · 被引用 1,493 次
- CLIPScore: A Reference-free Evaluation Metric for Image CaptioningJack Hessel, Ari Holtzman, Maxwell Forbes, Ronan Le Bras 等EMNLP 2021 · 被引用 937 次
相关 Paper
- One-Prompt-One-Story: Free-Lunch Consistent Text-to-Image Generation Using a Single PromptTao Liu, Kai Wang, Senmao Li, Joost van de Weijer 等ICLR 2025
- Infinite-Story: A Training-Free Consistent Text-to-Image GenerationJihun Park, Kyoungmin Lee, Jongmin Gim, Hyeonseo Jo 等AAAI 2026 · 被引用 1 次
- Training-Free Consistent Text-to-Image GenerationYoad Tewel, Omri Kaduri, Rinon Gal, Yoni Kasten 等SIGGRAPH 2024 · 被引用 57 次
- CoDi: Subject-Consistent and Pose-Diverse Text-to-Image GenerationZhanxin Gao, Beier Zhu, Liangyao, Jian Yang 等ICLR 2026 · 被引用 1 次
- StoryDiffusion: Consistent Self-Attention for Long-Range Image and Video GenerationYupeng Zhou, Daquan Zhou, Ming-Ming Cheng, Jiashi Feng 等NeurIPS 2024 · 被引用 291 次
