Similar Scenes Arouse Similar Emotions: Parallel Data Augmentation for Stylized Image Captioning
Guodun Li, Yuchen Zhai, Zehao Lin, Yin Zhang
摘要
Stylized image captioning systems aim to generate a caption not only semantically related to a given image but also consistent with a given style description. One of the biggest challenges with this task is the lack of sufficient paired stylized data. Many studies focus on unsupervised approaches, without considering from the perspective of data augmentation. We begin with the observation that people may recall similar emotions when they are in similar scenes, and often express similar emotions with similar style phrases, which underpins our data augmentation idea. In this paper, we propose a novel Extract-Retrieve-Generate data augmentation framework to extract style phrases from small-scale stylized sentences and graft them to large-scale factual captions. First, we design the emotional signal extractor to extract style phrases from small-scale stylized sentences. Second, we construct the plugable multi-modal scene retriever to retrieve scenes represented with pairs of an image and its stylized caption, which are similar to the query image or caption in the large-scale factual data. In the end, based on the style phrases of similar scenes and the factual description of the current scene, we build the emotion-aware caption generator to generate fluent and diversified stylized captions for the current scene. Extensive experimental results show that our framework can alleviate the data scarcity problem effectively. It also significantly boosts the performance of several existing image captioning models in both supervised and unsupervised settings, which outperforms the state-of-the-art stylized image captioning methods in terms of both sentence relevance and stylishness by a substantial margin.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Emotion-Prior Awareness Network for Emotional Video CaptioningPeipei Song, Dan Guo, Xun Yang, Shengeng Tang 等ACM MM 2023 · 被引用 29 次
- Impressions: Visual Semiotics and Aesthetic Impact UnderstandingJulia Kruk, Caleb Ziems, Diyi YangEMNLP 2023 · 被引用 5 次
- Visual Captioning at Will: Describing Images and Videos Guided by a Few Stylized SentencesDingyi Yang, Hongyu Chen, Xinglin Hou, Tiezheng Ge 等ACM MM 2023 · 被引用 5 次
- Attractive Storyteller: Stylized Visual Storytelling with Unpaired TextDingyi Yang, Qin JinACL 2023 · 被引用 1 次
它引用的顶会 Paper7
- BERT-ATTACK: Adversarial Attack Against BERT Using BERTLinyang Li, Ruotian Ma, Qipeng Guo, Xiangyang Xue 等EMNLP 2020 · 被引用 529 次
- Conditional Augmentation for Aspect Term Extraction via Masked Sequence-to-Sequence GenerationKun Li, Chengbo Chen, Xiaojun Quan, Qing Ling 等ACL 2020 · 被引用 101 次
- MemCap: Memorizing Style Knowledge for Image CaptioningWentian Zhao, Xinxiao Wu, Xiaoxun ZhangAAAI 2020 · 被引用 86 次
- Attention-Aware Polarity Sensitive Embedding for Affective Image RetrievalXingxu Yao, Dongyu She, Sicheng Zhao, Jie Liang 等ICCV 2019 · 被引用 31 次
- Self-Paced Video Data Augmentation by Generative Adversarial Networks with Insufficient SamplesYumeng Zhang, Gaoguo Jia, Li Chen, Mingrui Zhang 等ACM MM 2020 · 被引用 18 次
相关 Paper
- Detach and Attach: Stylized Image Captioning without Paired Stylized DatasetYutong Tan, Zheng Lin, Peng Fu, Mingyu Zheng 等ACM MM 2022 · 被引用 8 次
- UniStyle: Unified Style Modeling for Speaking Style Captioning and Stylistic Speech SynthesisXinfa Zhu, Wenjie Tian, Xinsheng Wang, Lei He 等ACM MM 2024 · 被引用 3 次
- Controllable Image Captioning via PromptingNing Wang, Jiahao Xie, Jihao Wu, Mingbo Jia 等AAAI 2023 · 被引用 43 次
- Paired Cross-Modal Data Augmentation for Fine-Grained Image-to-Text RetrievalHao Wang, Guosheng Lin, Steven C. H. Hoi, Chunyan MiaoACM MM 2022 · 被引用 12 次
- Unpaired Image Captioning via Scene Graph AlignmentsJiuxiang Gu, Shafiq R. Joty, Jianfei Cai, Handong Zhao 等ICCV 2019 · 被引用 191 次
