Attractive Storyteller: Stylized Visual Storytelling with Unpaired Text
Dingyi Yang, Qin Jin
Abstract
Most research on stylized image captioning aims to generate style-specific captions using unpaired text, and has achieved impressive performance for simple styles like positive and negative. However, unlike previous singlesentence captions whose style is mostly embodied in distinctive words or phrases, realworld styles are likely to be implied at the syntactic and discourse levels. In this work, we introduce a new task of Stylized Visual Storytelling (SVST), which aims to describe a photo stream with stylized stories that are more expressive and attractive. We propose a multitasking memory-augmented framework called StyleVSG, which is jointly trained on factual visual storytelling data and unpaired style corpus, achieving a trade-off between style accuracy and visual relevance. Particularly for unpaired stylized text, StyleVSG learns to reconstruct the stylistic story from roughly parallel visual inputs mined with the CLIP 1 model, avoiding problems caused by random mapping in previous methods. Furthermore, a memory module is designed to preserve the consistency and coherence of generated stories. Experiments show that our method can generate attractive and coherent stories with different styles, such as fairy tale, romance, and humor. The overall performance of our proposed StyleVSG surpasses state-of-the-art methods on both automatic and human evaluation metrics 2 .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 601f8b87-4d2f-48b0-8a20-518802a18fc5Cited by top-tier papers1
Ask how each one uses itBuilds on10
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- MART: Memory-Augmented Recurrent Transformer for Coherent Video Paragraph CaptioningJie Lei, Liwei Wang, Yelong Shen, Dong Yu et al.ACL 2020 · 168 citations
- MemCap: Memorizing Style Knowledge for Image CaptioningWentian Zhao, Xinxiao Wu, Xiaoxun ZhangAAAI 2020 · 86 citations
- Hooks in the Headline: Learning to Generate Headlines with Controlled StylesDi Jin, Zhijing Jin, Joey Tianyi Zhou, Lisa Orii et al.ACL 2020 · 56 citations
- Knowledge-Enriched Visual StorytellingChao-Chun Hsu, Zi-Yuan Chen, Chi-Yang Hsu, Chih-Chia Li et al.AAAI 2020 · 53 citations
Related papers
- Visual Captioning at Will: Describing Images and Videos Guided by a Few Stylized SentencesDingyi Yang, Hongyu Chen, Xinglin Hou, Tiezheng Ge et al.ACM MM 2023 · 5 citations
- Similar Scenes Arouse Similar Emotions: Parallel Data Augmentation for Stylized Image CaptioningGuodun Li, Yuchen Zhai, Zehao Lin, Yin ZhangACM MM 2021 · 23 citations
- Detach and Attach: Stylized Image Captioning without Paired Stylized DatasetYutong Tan, Zheng Lin, Peng Fu, Mingyu Zheng et al.ACM MM 2022 · 8 citations
- Hide-and-Tell: Learning to Bridge Photo Streams for Visual StorytellingYunjae Jung, Dahun Kim, Sanghyun Woo, Kyungsu Kim et al.AAAI 2020 · 35 citations
- Imagine, Reason and Write: Visual Storytelling with Graph Knowledge and Relational ReasoningChunpu Xu, Min Yang, Chengming Li, Ying Shen et al.AAAI 2021 · 39 citations
