Detach and Attach: Stylized Image Captioning without Paired Stylized Dataset
Yutong Tan, Zheng Lin, Peng Fu, Mingyu Zheng, Lanrui Wang, Yanan Cao, Weiping Wang
摘要
Stylized Image Captioning aims to generate captions with accurate image content and stylized elements simultaneously. However, large-scaled image and stylized caption pairs cost lots of resources and are usually unavailable. Therefore, it's a challenge to generate stylized captions without paired stylized caption dataset. Previous work on controlling the style of generated captions in an unsupervised way can be divided into two ways: implicitly and explicitly. The former mainly relies on a well-trained language model to capture style knowledge, which is limited to a single style and hard to handle multi-style task. Thus, the latter uses extra style constraints such as outlined style labels or stylized words extracted from stylized sentences to control the style rather than the trained style-specific language model. However, certain styles, such as humorous and romance, are implied in the whole sentence, instead of in some words of a sentence. To address the problems above, we propose a two-step method based on Transformer: firstly detach style representations from large-scaled stylized text-only corpus to provide more holistic style supervision, and secondly attach the style representations to image content to generate stylized captions. We learn a shared image-text space to narrow the gap between the image and the text modality for better attachment. Due to the trade-off between semantics and style, we explore three injection methods of style representations to balance two requirements of image content preservation and stylization. Experiments show that our method outperforms the state-of-the-art systems in overall performance, especially on implied styles.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
引用它的顶会 Paper5
- I can't believe there's no images! : Learning Visual Tasks Using Only Language SupervisionSophia Gu, Christopher Clark, Aniruddha KembhaviICCV 2023 · 被引用 41 次
- Impressions: Visual Semiotics and Aesthetic Impact UnderstandingJulia Kruk, Caleb Ziems, Diyi YangEMNLP 2023 · 被引用 5 次
- Visual Captioning at Will: Describing Images and Videos Guided by a Few Stylized SentencesDingyi Yang, Hongyu Chen, Xinglin Hou, Tiezheng Ge 等ACM MM 2023 · 被引用 5 次
- BiMa: Towards Biases Mitigation for Text-Video Retrieval via Scene Element GuidanceHuy Le, Nhat Chung, Tung Kieu, Anh Nguyen 等ACM MM 2025 · 被引用 2 次
- Attractive Storyteller: Stylized Visual Storytelling with Unpaired TextDingyi Yang, Qin JinACL 2023 · 被引用 1 次
相关 Paper
- MemCap: Memorizing Style Knowledge for Image CaptioningWentian Zhao, Xinxiao Wu, Xiaoxun ZhangAAAI 2020 · 被引用 86 次
- Similar Scenes Arouse Similar Emotions: Parallel Data Augmentation for Stylized Image CaptioningGuodun Li, Yuchen Zhai, Zehao Lin, Yin ZhangACM MM 2021 · 被引用 23 次
- Towards Unsupervised Image Captioning With Shared Multimodal EmbeddingsIro Laina, Christian Rupprecht, Nassir NavabICCV 2019 · 被引用 115 次
- UniStyle: Unified Style Modeling for Speaking Style Captioning and Stylistic Speech SynthesisXinfa Zhu, Wenjie Tian, Xinsheng Wang, Lei He 等ACM MM 2024 · 被引用 3 次
- Controllable Image Captioning via PromptingNing Wang, Jiahao Xie, Jihao Wu, Mingbo Jia 等AAAI 2023 · 被引用 43 次
