Character-centric Story Visualization via Visual Planning and Token Alignment
Hong Chen, Rujun Han, Te-Lin Wu, Hideki Nakayama, Nanyun Peng
Abstract
Story visualization advances the traditional text-to-image generation by enabling multiple image generation based on a complete story. This task requires machines to 1) understand long text inputs and 2) produce a globally consistent image sequence that illustrates the contents of the story. A key challenge of consistent story visualization is to preserve characters that are essential in stories. To tackle the challenge, we propose to adapt a recent work that augments Vector-Quantized Variational Autoencoders (VQ-VAE) with a text-tovisual-token (transformer) architecture. Specifically, we modify the text-to-visual-token module with a two-stage framework: 1) character token planning model that predicts the visual tokens for characters only; 2) visual token completion model that generates the remaining visual token sequence, which is sent to VQ-VAE for finalizing image generations. To encourage characters to appear in the images, we further train the two-stage framework with a character-token alignment objective. Extensive experiments and evaluations demonstrate that the proposed method excels at preserving characters and can produce higher quality image sequences compared with the strong baselines. Code can be found in https: //github.com/PlusLabNLP/VP-CSV
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers11
- Boosting Consistency in Story Visualization with Rich-Contextual Conditional Diffusion ModelsFei Shen, Hu Ye, Sibo Liu, Jun Zhang et al.AAAI 2025 · 74 citations
- Controllable Text Generation with Neurally-Decomposed OracleTao Meng, Sidi Lu, Nanyun Peng, Kai-Wei ChangNeurIPS 2022 · 45 citations
- ViStoryBench: Comprehensive Benchmark Suite for Story VisualizationCailin Zhuang, Ailin Huang, Hu Yaoqi, Jingwei Wu et al.CVPR 2026 · 37 citations
- Intelligent Grimm - Open-ended Visual Storytelling via Latent Diffusion ModelsChang Liu, Haoning Wu, Yujie Zhong, Xiaoyun Zhang et al.CVPR 2024 · 32 citations
- Story-Iter: A Training-free Iterative Paradigm for Long Story VisualizationJiawei Mao, Xiaoke Huang, Yunfei Xie, Yuanqi Chang et al.ICLR 2026 · 18 citations
Builds on4
- Zero-Shot Text-to-Image GenerationAditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray et al.ICML 2021 · 6,356 citations
- CogView: Mastering Text-to-Image Generation via TransformersMing Ding, Zhuoyi Yang, Wenyi Hong, Wendi Zheng et al.NeurIPS 2021 · 1,026 citations
- Content Planning for Neural Story Generation with Aristotelian RescoringSeraphina Goldfarb-Tarrant, Tuhin Chakrabarty, Ralph M. Weischedel, Nanyun PengEMNLP 2020 · 106 citations
- Integrating Visuospatial, Linguistic, and Commonsense Structure into Story VisualizationAdyasha Maharana, Mohit BansalEMNLP 2021 · 37 citations
Related papers
- Story Visualization by Online Text Augmentation with Context MemoryDaechul Ahn, Daneul Kim, Gwangmo Song, Seung Hwan Kim et al.ICCV 2023 · 11 citations
- Towards Accurate Image Coding: Improved Autoregressive Image Generation with Dynamic Vector QuantizationMengqi Huang, Zhendong Mao, Zhuowei Chen, Yongdong ZhangCVPR 2023
- A Character-Centric Neural Model for Automated Story GenerationDanyang Liu, Juntao Li, Meng-Hsuan Yu, Ziming Huang et al.AAAI 2020 · 47 citations
- VQRAE: Representation Quantization Autoencoders for Multimodal Understanding, Generation and ReconstructionSinan Du, Jiahao Guo, Bo Li, Shuhao Cui et al.CVPR 2026 · 11 citations
- ContextualStory: Consistent Visual Storytelling with Spatially-Enhanced and Storyline ContextSixiao Zheng, Yanwei FuAAAI 2025 · 12 citations
