Story Visualization by Online Text Augmentation with Context Memory
Daechul Ahn, Daneul Kim, Gwangmo Song, Seung Hwan Kim, Honglak Lee, Dongyeop Kang, Jonghyun Choi
Abstract
Story visualization (SV) is a challenging text-to-image generation task for the difficulty of not only rendering visual details from the text descriptions but also encoding a long-term context across multiple sentences. While prior efforts mostly focus on generating a semantically relevant image for each sentence, encoding a context spread across the given paragraph to generate contextually convincing images (e.g., with a correct character or with a proper background of the scene) remains a challenge. To this end, we propose a novel memory architecture for the Bi-directional Transformer framework with an online text augmentation that generates multiple pseudo-descriptions as supplementary supervision during training for better generalization to the language variation at inference. In extensive experiments on the two popular SV benchmarks, i.e., the Pororo-SV and Flintstones-SV, the proposed method significantly outperforms the state of the arts in various metrics including FID, character F1, frame accuracy, BLEU-2/3, and R-precision with similar or less computational complexity.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext fe2fc2b6-c1ba-428c-8321-9d51f533ebbbCited by top-tier papers5
- Boosting Consistency in Story Visualization with Rich-Contextual Conditional Diffusion ModelsFei Shen, Hu Ye, Sibo Liu, Jun Zhang et al.AAAI 2025 · 74 citations
- Story-Iter: A Training-free Iterative Paradigm for Long Story VisualizationJiawei Mao, Xiaoke Huang, Yunfei Xie, Yuanqi Chang et al.ICLR 2026 · 18 citations
- ContextualStory: Consistent Visual Storytelling with Spatially-Enhanced and Storyline ContextSixiao Zheng, Yanwei FuAAAI 2025 · 12 citations
- IP-Prompter: Training-Free Theme-Specific Image Generation via Dynamic Visual PromptingYuxin Zhang, Minyan Luo, Weiming Dong, Xiao Yang et al.SIGGRAPH 2025 · 2 citations
- VisAug: Facilitating Speech-Rich Web Video Navigation and Engagement with Auto-Generated Visual AugmentationsBaoquan Zhao, Xiaofan Ma, Qianshi Pang, Ruomei Wang et al.ACM MM 2025 · 1 citation
Builds on13
- Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingChitwan Saharia, William Chan, Saurabh Saxena, Lala Li et al.NeurIPS 2022 · 8,965 citations
- BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and GenerationJunnan Li, Dongxu Li, Caiming Xiong, Steven C. H. HoiICML 2022 · 6,549 citations
- Zero-Shot Text-to-Image GenerationAditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray et al.ICML 2021 · 6,356 citations
- Generative Pretraining From PixelsMark Chen, Alec Radford, Rewon Child, Jeffrey Wu et al.ICML 2020 · 1,773 citations
- Contrastive Representation DistillationYonglong Tian, Dilip Krishnan, Phillip IsolaICLR 2020 · 1,305 citations
Related papers
- Character-centric Story Visualization via Visual Planning and Token AlignmentHong Chen, Rujun Han, Te-Lin Wu, Hideki Nakayama et al.EMNLP 2022 · 18 citations
- MMT: Image-guided Story Ending Generation with Multimodal Memory TransformerDizhan Xue, Shengsheng Qian, Quan Fang, Changsheng XuACM MM 2022 · 16 citations
- Latent Memory-augmented Graph Transformer for Visual StorytellingMengshi Qi, Jie Qin, Di Huang, Zhiqiang Shen et al.ACM MM 2021 · 18 citations
- Clustering Generative Adversarial Networks for Story VisualizationBowen Li, Philip H. S. Torr, Thomas LukasiewiczACM MM 2022 · 6 citations
- Show Me What and Tell Me How: Video Synthesis via Multimodal ConditioningLigong Han, Jian Ren, Hsin-Ying Lee, Francesco Barbieri et al.CVPR 2022 · 36 citations
