Integrating Visuospatial, Linguistic, and Commonsense Structure into Story Visualization
Adyasha Maharana, Mohit Bansal
摘要
While much research has been done in textto-image synthesis, little work has been done to explore the usage of linguistic structure of the input text. Such information is even more important for story visualization since its inputs have an explicit narrative structure that needs to be translated into an image sequence (or visual story). Prior work in this domain has shown that there is ample room for improvement in the generated image sequence in terms of visual quality, consistency and relevance. In this paper, we first explore the use of constituency parse trees using a Transformer-based recurrent architecture for encoding structured input. Second, we augment the structured input with commonsense information and study the impact of this external knowledge on the generation of visual story. Third, we also incorporate visual structure via bounding boxes and dense captioning to provide feedback about the characters/objects in generated images within a dual learning setup. We show that off-theshelf dense-captioning models trained on Visual Genome can improve the spatial structure of images from a different target domain without needing fine-tuning. We train the model end-to-end using intra-story contrastive loss (between words and image sub-regions) and show significant improvements in visual quality. Finally, we provide an analysis of the linguistic and visuo-spatial information. 1 1 Code and data: https://github.com/ adymaharana/VLCStoryGan . Ground Truth DuCo Captions VLC Petty asks whether it is because of cookies. Eddy denies with his hands. Petty hands her cookies to Eddy. Petty gives her cookies to Loopy and Crong. Crong sighs.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper18
- Boosting Consistency in Story Visualization with Rich-Contextual Conditional Diffusion ModelsFei Shen, Hu Ye, Sibo Liu, Jun Zhang 等AAAI 2025 · 被引用 74 次
- ViStoryBench: Comprehensive Benchmark Suite for Story VisualizationCailin Zhuang, Ailin Huang, Hu Yaoqi, Jingwei Wu 等CVPR 2026 · 被引用 37 次
- Intelligent Grimm - Open-ended Visual Storytelling via Latent Diffusion ModelsChang Liu, Haoning Wu, Yujie Zhong, Xiaoyun Zhang 等CVPR 2024 · 被引用 32 次
- Character-centric Story Visualization via Visual Planning and Token AlignmentHong Chen, Rujun Han, Te-Lin Wu, Hideki Nakayama 等EMNLP 2022 · 被引用 18 次
- Story-Iter: A Training-free Iterative Paradigm for Long Story VisualizationJiawei Mao, Xiaoke Huang, Yunfei Xie, Yuanqi Chang 等ICLR 2026 · 被引用 18 次
它引用的顶会 Paper7
- Zero-Shot Text-to-Image GenerationAditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray 等ICML 2021 · 被引用 6,356 次
- ContraGAN: Contrastive Learning for Conditional Image GenerationMinguk Kang, Jaesik ParkNeurIPS 2020 · 被引用 216 次
- MART: Memory-Augmented Recurrent Transformer for Coherent Video Paragraph CaptioningJie Lei, Liwei Wang, Yelong Shen, Dong Yu 等ACL 2020 · 被引用 168 次
- X-LXMERT: Paint, Caption and Answer Questions with Multi-Modal TransformersJaemin Cho, Jiasen Lu, Dustin Schwenk, Hannaneh Hajishirzi 等EMNLP 2020 · 被引用 80 次
- Tree-Structured Attention with Hierarchical AccumulationXuan-Phi Nguyen, Shafiq R. Joty, Steven C. H. Hoi, Richard SocherICLR 2020 · 被引用 79 次
相关 Paper
- Pixel Aligned Language ModelsJiarui Xu, Xingyi Zhou, Shen Yan, Xiuye Gu 等CVPR 2024 · 被引用 6 次
- Text-Only Training for Visual StorytellingYuechen Wang, Wengang Zhou, Zhenbo Lu, Houqiang LiACM MM 2023 · 被引用 4 次
- UniT3D: A Unified Transformer for 3D Dense Captioning and Visual GroundingDave Zhenyu Chen, Ronghang Hu, Xinlei Chen, Matthias Nießner 等ICCV 2023 · 被引用 82 次
- Imagine, Reason and Write: Visual Storytelling with Graph Knowledge and Relational ReasoningChunpu Xu, Min Yang, Chengming Li, Ying Shen 等AAAI 2021 · 被引用 39 次
- Auto-Parsing Network for Image Captioning and Visual Question AnsweringXu Yang, Chongyang Gao, Hanwang Zhang, Jianfei CaiICCV 2021 · 被引用 45 次
