Detach and Attach: Stylized Image Captioning without Paired Stylized Dataset
Yutong Tan, Zheng Lin, Peng Fu, Mingyu Zheng, Lanrui Wang, Yanan Cao, Weiping Wang
Abstract
Stylized Image Captioning aims to generate captions with accurate image content and stylized elements simultaneously. However, large-scaled image and stylized caption pairs cost lots of resources and are usually unavailable. Therefore, it's a challenge to generate stylized captions without paired stylized caption dataset. Previous work on controlling the style of generated captions in an unsupervised way can be divided into two ways: implicitly and explicitly. The former mainly relies on a well-trained language model to capture style knowledge, which is limited to a single style and hard to handle multi-style task. Thus, the latter uses extra style constraints such as outlined style labels or stylized words extracted from stylized sentences to control the style rather than the trained style-specific language model. However, certain styles, such as humorous and romance, are implied in the whole sentence, instead of in some words of a sentence. To address the problems above, we propose a two-step method based on Transformer: firstly detach style representations from large-scaled stylized text-only corpus to provide more holistic style supervision, and secondly attach the style representations to image content to generate stylized captions. We learn a shared image-text space to narrow the gap between the image and the text modality for better attachment. Due to the trade-off between semantics and style, we explore three injection methods of style representations to balance two requirements of image content preservation and stylization. Experiments show that our method outperforms the state-of-the-art systems in overall performance, especially on implied styles.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 00591cf5-ee71-45ee-b923-052b860ed558Cited by top-tier papers5
- I can't believe there's no images! : Learning Visual Tasks Using Only Language SupervisionSophia Gu, Christopher Clark, Aniruddha KembhaviICCV 2023 · 41 citations
- Impressions: Visual Semiotics and Aesthetic Impact UnderstandingJulia Kruk, Caleb Ziems, Diyi YangEMNLP 2023 · 5 citations
- Visual Captioning at Will: Describing Images and Videos Guided by a Few Stylized SentencesDingyi Yang, Hongyu Chen, Xinglin Hou, Tiezheng Ge et al.ACM MM 2023 · 5 citations
- BiMa: Towards Biases Mitigation for Text-Video Retrieval via Scene Element GuidanceHuy Le, Nhat Chung, Tung Kieu, Anh Nguyen et al.ACM MM 2025 · 2 citations
- Attractive Storyteller: Stylized Visual Storytelling with Unpaired TextDingyi Yang, Qin JinACL 2023 · 1 citation
Related papers
- MemCap: Memorizing Style Knowledge for Image CaptioningWentian Zhao, Xinxiao Wu, Xiaoxun ZhangAAAI 2020 · 86 citations
- Similar Scenes Arouse Similar Emotions: Parallel Data Augmentation for Stylized Image CaptioningGuodun Li, Yuchen Zhai, Zehao Lin, Yin ZhangACM MM 2021 · 23 citations
- Towards Unsupervised Image Captioning With Shared Multimodal EmbeddingsIro Laina, Christian Rupprecht, Nassir NavabICCV 2019 · 115 citations
- UniStyle: Unified Style Modeling for Speaking Style Captioning and Stylistic Speech SynthesisXinfa Zhu, Wenjie Tian, Xinsheng Wang, Lei He et al.ACM MM 2024 · 3 citations
- Controllable Image Captioning via PromptingNing Wang, Jiahao Xie, Jihao Wu, Mingbo Jia et al.AAAI 2023 · 43 citations
