Clustering Generative Adversarial Networks for Story Visualization
Bowen Li, Philip H. S. Torr, Thomas Lukasiewicz
Abstract
Story visualization aims to generate a series of images, semantically matching a given sequence of sentences, one for each, and different output images within a story should be consistent with each other. Current methods generate story images by using a heavy architecture with two generative adversarial networks (GANs), one for image quality, and one for story consistency, and also rely on additional segmentation masks or auxiliary captioning networks. In this paper, we aim to build a concise and single-GAN-based network, neither depending on additional semantic information nor captioning networks. To achieve this, we propose a contrastive-learning- and clustering-learning-based approach for story visualization. Our network utilizes contrastive losses between language and visual information to maximize the mutual information between them, and further extends it with clustering learning in the training process to capture semantic similarity across modalities. So, the discriminator in our approach provides comprehensive feedback to the generator, regarding both image quality and story consistency at the same time, allowing to have a single-GAN-based network to produce high-quality synthetic results. Extensive experiments on two datasets demonstrate that our single-GAN-based network has a smaller number of total parameters in the network, but achieves a major step up from previous methods, which improves FID from 78.64 to 39.17, and FSD from 94.53 to 41.18 on Pororo-SV, and establishes a strong benchmark FID of 76.51 and FSD of 19.74 on Abstract Scenes.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext ae94e32d-791c-48fa-8c98-85d25caab480Cited by top-tier papers3
- ContextualStory: Consistent Visual Storytelling with Spatially-Enhanced and Storyline ContextSixiao Zheng, Yanwei FuAAAI 2025 · 12 citations
- StoryWeaver: A Unified World Model for Knowledge-Enhanced Story Character CustomizationJinlu Zhang, Jiji Tang, Rongsheng Zhang, Tangjie Lv et al.AAAI 2025 · 3 citations
- Integrating Sequence and Image Modeling in Irregular Medical Time Series Through Self-Supervised LearningLiuqing Chen, Shuhong Xiao, Shixian Ding, Shanhai Hu et al.AAAI 2025 · 3 citations
Builds on16
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- Zero-Shot Text-to-Image GenerationAditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray et al.ICML 2021 · 6,356 citations
- Unsupervised Learning of Visual Features by Contrasting Cluster AssignmentsMathilde Caron, Ishan Misra, Julien Mairal, Priya Goyal et al.NeurIPS 2020 · 5,249 citations
- GLIDE: Towards Photorealistic Image Generation and Editing with Text-Guided Diffusion ModelsAlexander Quinn Nichol, Prafulla Dhariwal, Aditya Ramesh, Pranav Shyam et al.ICML 2022 · 4,691 citations
- Self-labelling via simultaneous clustering and representation learningYuki Markus Asano, Christian Rupprecht, Andrea VedaldiICLR 2020 · 873 citations
Related papers
- Story Visualization by Online Text Augmentation with Context MemoryDaechul Ahn, Daneul Kim, Gwangmo Song, Seung Hwan Kim et al.ICCV 2023 · 11 citations
- Cross-Modal Contrastive Learning for Text-to-Image GenerationHan Zhang, Jing Yu Koh, Jason Baldridge, Honglak Lee et al.CVPR 2021
- Learning Visual Representations via Language-Guided SamplingMohamed El Banani, Karan Desai, Justin JohnsonCVPR 2023
- Text-Only Training for Visual StorytellingYuechen Wang, Wengang Zhou, Zhenbo Lu, Houqiang LiACM MM 2023 · 4 citations
- SSCL: Adversarially Guided Image Compression via Semantic and Spectral Consistency LearningWei Jiang, Yongqi Zhai, Jiayu Yang, Bohao Feng et al.AAAI 2026
