Topic Adaptation and Prototype Encoding for Few-Shot Visual Storytelling
Jiacheng Li, Siliang Tang, Juncheng Li, Jun Xiao, Fei Wu, Shiliang Pu, Yueting Zhuang
摘要
Visual Storytelling (VIST) is a task to tell a narrative story about a certain topic according to the given photo stream. The existing studies focus on designing complex models, which rely on a huge amount of human-annotated data. However, the annotation of VIST is extremely costly and many topics cannot be covered in the training dataset due to the long-tail topic distribution. In this paper, we focus on enhancing the generalization ability of the VIST model by considering the few-shot setting. Inspired by the way humans tell a story, we propose a topic adaptive storyteller to model the ability of inter-topic generalization. In practice, we apply the gradient-based meta-learning algorithm on multi-modal seq2seq models to endow the model the ability to adapt quickly from topic to topic. Besides, We further propose a prototype encoding structure to model the ability of intra-topic derivation. Specifically, we encode and restore the few training story text to serve as a reference to guide the generation at inference time. Experimental results show that topic adaptation and prototype encoding structure mutually bring benefit to the few-shot model on BLEU and METEOR metric. The further case study shows that the stories generated after few-shot adaptation are more relative and expressive.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Gradient-Regulated Meta-Prompt Learning for Generalizable Vision-Language ModelsJuncheng Li, Minghe Gao, Longhui Wei, Siliang Tang 等ICCV 2023 · 被引用 34 次
- Latent Memory-augmented Graph Transformer for Visual StorytellingMengshi Qi, Jie Qin, Di Huang, Zhiqiang Shen 等ACM MM 2021 · 被引用 18 次
- Dilated Context Integrated Network with Cross-Modal Consensus for Temporal Emotion Localization in VideosJuncheng Li, Junlin Xie, Linchao Zhu, Long Qian 等ACM MM 2022 · 被引用 8 次
- Multi-Modal Experience Inspired AI CreationQian Cao, Xu Chen, Ruihua Song, Hao Jiang 等ACM MM 2022 · 被引用 3 次
它引用的顶会 Paper3
- What Makes A Good Story? Designing Composite Rewards for Visual StorytellingJunjie Hu, Yu Cheng, Zhe Gan, Jingjing Liu 等AAAI 2020 · 被引用 73 次
- Knowledge-Enriched Visual StorytellingChao-Chun Hsu, Zi-Yuan Chen, Chi-Yang Hsu, Chih-Chia Li 等AAAI 2020 · 被引用 53 次
- Hide-and-Tell: Learning to Bridge Photo Streams for Visual StorytellingYunjae Jung, Dahun Kim, Sanghyun Woo, Kyungsu Kim 等AAAI 2020 · 被引用 35 次
相关 Paper
- OneStory: Coherent Multi-Shot Video Generation with Adaptive MemoryZhaochong An, Menglin Jia, Haonan Qiu, Zijian Zhou 等CVPR 2026 · 被引用 33 次
- Meta Learning to Bridge Vision and Language Models for Multimodal Few-Shot LearningIvona Najdenkoska, Xiantong Zhen, Marcel WorringICLR 2023 · 被引用 8 次
- Text-Only Training for Visual StorytellingYuechen Wang, Wengang Zhou, Zhenbo Lu, Houqiang LiACM MM 2023 · 被引用 4 次
- Learning to Learn Variational Semantic MemoryXiantong Zhen, Ying-Jun Du, Huan Xiong, Qiang Qiu 等NeurIPS 2020 · 被引用 40 次
- Adaptive FSS: A Novel Few-Shot Segmentation Framework via Prototype EnhancementJing Wang, Jiangyun Li, Chen Chen, Yisi Zhang 等AAAI 2024 · 被引用 24 次
