StoryTailor: A Zero-Shot Pipeline for Action-Rich Multi-Subject Visual Narratives
Jinghao Hu, Yuhe Zhang, Guohua Geng, Kang Li, Han Zhang
Abstract
Generating multi-frame, action-rich visual narratives without fine-tuning faces a threefold tension: action text faithfulness, subject identity fidelity, and cross frame background continuity. We propose StoryTailor, a zero-shot pipeline that runs on a single RTX 4090 (24 GB) and produces temporally coherent, identity-preserving image sequences from a long narrative prompt, per subject references, and grounding boxes. Three synergistic modules drive the system: Gaussian-Centered Attention (GCA) to dynamically focus on each subject core and ease groundingbox overlaps; Action-Boost Singular Value Reweighting (AB-SVR) to amplify action-related directions in the text embedding space; and Selective Forgetting Cache (SFC) that retains transferable background cues, forgets nonessential history, and selectively surfaces the retained cues to build cross scene semantic ties. Compared with baseline methods, the experiments show that CLIP-T improves by up to 10-15%, with DreamSim lower than strong baselines, while CLIP-I stays in a visually acceptable, competitive range. With a matched resolution and steps on a 24 GB GPU, inference is faster than FluxKontext. Qualitatively, StoryTailor delivers expressive interactions and evolving yet stable scenes.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on23
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- SDXL: Improving Latent Diffusion Models for High-Resolution Image SynthesisDustin Podell, Zion English, Kyle Lacey, Andreas Blattmann et al.ICLR 2024 · 4,569 citations
Related papers
- TS-Attn: Temporal-wise Separable Attention for Multi-Event Video GenerationHongyu Zhang, Yufan Deng, Zilin Pan, Peng-Tao Jiang et al.ICLR 2026 · 5 citations
- DreamRunner: Fine-Grained Compositional Story-to-Video Generation with Retrieval-Augmented Motion AdaptationZun Wang, Jialu Li, Han Lin, Jaehong Yoon et al.AAAI 2026 · 5 citations
- SSR-Encoder: Encoding Selective Subject Representation for Subject-Driven GenerationYuxuan Zhang, Yiren Song, Jiaming Liu, Rui Wang et al.CVPR 2024 · 34 citations
- OneStory: Coherent Multi-Shot Video Generation with Adaptive MemoryZhaochong An, Menglin Jia, Haonan Qiu, Zijian Zhou et al.CVPR 2026 · 33 citations
- Infinite-Story: A Training-Free Consistent Text-to-Image GenerationJihun Park, Kyoungmin Lee, Jongmin Gim, Hyeonseo Jo et al.AAAI 2026 · 1 citation
