Integrating Visuospatial, Linguistic, and Commonsense Structure into Story Visualization
Adyasha Maharana, Mohit Bansal
Abstract
While much research has been done in textto-image synthesis, little work has been done to explore the usage of linguistic structure of the input text. Such information is even more important for story visualization since its inputs have an explicit narrative structure that needs to be translated into an image sequence (or visual story). Prior work in this domain has shown that there is ample room for improvement in the generated image sequence in terms of visual quality, consistency and relevance. In this paper, we first explore the use of constituency parse trees using a Transformer-based recurrent architecture for encoding structured input. Second, we augment the structured input with commonsense information and study the impact of this external knowledge on the generation of visual story. Third, we also incorporate visual structure via bounding boxes and dense captioning to provide feedback about the characters/objects in generated images within a dual learning setup. We show that off-theshelf dense-captioning models trained on Visual Genome can improve the spatial structure of images from a different target domain without needing fine-tuning. We train the model end-to-end using intra-story contrastive loss (between words and image sub-regions) and show significant improvements in visual quality. Finally, we provide an analysis of the linguistic and visuo-spatial information. 1 1 Code and data: https://github.com/ adymaharana/VLCStoryGan . Ground Truth DuCo Captions VLC Petty asks whether it is because of cookies. Eddy denies with his hands. Petty hands her cookies to Eddy. Petty gives her cookies to Loopy and Crong. Crong sighs.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 50dcd6a9-2de9-4f7e-84cf-1792fa54fbd5Cited by top-tier papers18
- Boosting Consistency in Story Visualization with Rich-Contextual Conditional Diffusion ModelsFei Shen, Hu Ye, Sibo Liu, Jun Zhang et al.AAAI 2025 · 74 citations
- ViStoryBench: Comprehensive Benchmark Suite for Story VisualizationCailin Zhuang, Ailin Huang, Hu Yaoqi, Jingwei Wu et al.CVPR 2026 · 37 citations
- Intelligent Grimm - Open-ended Visual Storytelling via Latent Diffusion ModelsChang Liu, Haoning Wu, Yujie Zhong, Xiaoyun Zhang et al.CVPR 2024 · 32 citations
- Character-centric Story Visualization via Visual Planning and Token AlignmentHong Chen, Rujun Han, Te-Lin Wu, Hideki Nakayama et al.EMNLP 2022 · 18 citations
- Story-Iter: A Training-free Iterative Paradigm for Long Story VisualizationJiawei Mao, Xiaoke Huang, Yunfei Xie, Yuanqi Chang et al.ICLR 2026 · 18 citations
Builds on7
- Zero-Shot Text-to-Image GenerationAditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray et al.ICML 2021 · 6,356 citations
- ContraGAN: Contrastive Learning for Conditional Image GenerationMinguk Kang, Jaesik ParkNeurIPS 2020 · 216 citations
- MART: Memory-Augmented Recurrent Transformer for Coherent Video Paragraph CaptioningJie Lei, Liwei Wang, Yelong Shen, Dong Yu et al.ACL 2020 · 168 citations
- X-LXMERT: Paint, Caption and Answer Questions with Multi-Modal TransformersJaemin Cho, Jiasen Lu, Dustin Schwenk, Hannaneh Hajishirzi et al.EMNLP 2020 · 80 citations
- Tree-Structured Attention with Hierarchical AccumulationXuan-Phi Nguyen, Shafiq R. Joty, Steven C. H. Hoi, Richard SocherICLR 2020 · 79 citations
Related papers
- Pixel Aligned Language ModelsJiarui Xu, Xingyi Zhou, Shen Yan, Xiuye Gu et al.CVPR 2024 · 6 citations
- Text-Only Training for Visual StorytellingYuechen Wang, Wengang Zhou, Zhenbo Lu, Houqiang LiACM MM 2023 · 4 citations
- UniT3D: A Unified Transformer for 3D Dense Captioning and Visual GroundingDave Zhenyu Chen, Ronghang Hu, Xinlei Chen, Matthias Nießner et al.ICCV 2023 · 82 citations
- Imagine, Reason and Write: Visual Storytelling with Graph Knowledge and Relational ReasoningChunpu Xu, Min Yang, Chengming Li, Ying Shen et al.AAAI 2021 · 39 citations
- Auto-Parsing Network for Image Captioning and Visual Question AnsweringXu Yang, Chongyang Gao, Hanwang Zhang, Jianfei CaiICCV 2021 · 45 citations
