Topic Adaptation and Prototype Encoding for Few-Shot Visual Storytelling
Jiacheng Li, Siliang Tang, Juncheng Li, Jun Xiao, Fei Wu, Shiliang Pu, Yueting Zhuang
Abstract
Visual Storytelling (VIST) is a task to tell a narrative story about a certain topic according to the given photo stream. The existing studies focus on designing complex models, which rely on a huge amount of human-annotated data. However, the annotation of VIST is extremely costly and many topics cannot be covered in the training dataset due to the long-tail topic distribution. In this paper, we focus on enhancing the generalization ability of the VIST model by considering the few-shot setting. Inspired by the way humans tell a story, we propose a topic adaptive storyteller to model the ability of inter-topic generalization. In practice, we apply the gradient-based meta-learning algorithm on multi-modal seq2seq models to endow the model the ability to adapt quickly from topic to topic. Besides, We further propose a prototype encoding structure to model the ability of intra-topic derivation. Specifically, we encode and restore the few training story text to serve as a reference to guide the generation at inference time. Experimental results show that topic adaptation and prototype encoding structure mutually bring benefit to the few-shot model on BLEU and METEOR metric. The further case study shows that the stories generated after few-shot adaptation are more relative and expressive.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 5cf0838b-c34a-4087-b093-6cb3a0d95322Cited by top-tier papers4
- Gradient-Regulated Meta-Prompt Learning for Generalizable Vision-Language ModelsJuncheng Li, Minghe Gao, Longhui Wei, Siliang Tang et al.ICCV 2023 · 34 citations
- Latent Memory-augmented Graph Transformer for Visual StorytellingMengshi Qi, Jie Qin, Di Huang, Zhiqiang Shen et al.ACM MM 2021 · 18 citations
- Dilated Context Integrated Network with Cross-Modal Consensus for Temporal Emotion Localization in VideosJuncheng Li, Junlin Xie, Linchao Zhu, Long Qian et al.ACM MM 2022 · 8 citations
- Multi-Modal Experience Inspired AI CreationQian Cao, Xu Chen, Ruihua Song, Hao Jiang et al.ACM MM 2022 · 3 citations
Builds on3
- What Makes A Good Story? Designing Composite Rewards for Visual StorytellingJunjie Hu, Yu Cheng, Zhe Gan, Jingjing Liu et al.AAAI 2020 · 73 citations
- Knowledge-Enriched Visual StorytellingChao-Chun Hsu, Zi-Yuan Chen, Chi-Yang Hsu, Chih-Chia Li et al.AAAI 2020 · 53 citations
- Hide-and-Tell: Learning to Bridge Photo Streams for Visual StorytellingYunjae Jung, Dahun Kim, Sanghyun Woo, Kyungsu Kim et al.AAAI 2020 · 35 citations
Related papers
- OneStory: Coherent Multi-Shot Video Generation with Adaptive MemoryZhaochong An, Menglin Jia, Haonan Qiu, Zijian Zhou et al.CVPR 2026 · 33 citations
- Meta Learning to Bridge Vision and Language Models for Multimodal Few-Shot LearningIvona Najdenkoska, Xiantong Zhen, Marcel WorringICLR 2023 · 8 citations
- Text-Only Training for Visual StorytellingYuechen Wang, Wengang Zhou, Zhenbo Lu, Houqiang LiACM MM 2023 · 4 citations
- Learning to Learn Variational Semantic MemoryXiantong Zhen, Ying-Jun Du, Huan Xiong, Qiang Qiu et al.NeurIPS 2020 · 40 citations
- Adaptive FSS: A Novel Few-Shot Segmentation Framework via Prototype EnhancementJing Wang, Jiangyun Li, Chen Chen, Yisi Zhang et al.AAAI 2024 · 24 citations
