Image Change Captioning by Learning From an Auxiliary Task
Mehrdad Hosseinzadeh, Yang Wang
Abstract
We tackle the challenging task of image change captioning. The goal is to describe the subtle difference between two very similar images by generating a sentence caption. While the recent methods mainly focus on proposing new model architectures for this problem, we instead focus on an alternative training scheme. Inspired by the success of multi-task learning, we formulate a training scheme that uses an auxiliary task to improve the training of the change captioning network. We argue that the task of composed query image retrieval is a natural choice as the auxiliary task. Given two almost similar images as the input, the primary network generates a caption describing the fine change between those two images. Next, the auxiliary network is provided with the generated caption and one of those two images. It then tries to pick the second image among a set of candidates. This forces the primary network to generate detailed and precise captions via having an extra supervision loss by the auxiliary network. Furthermore, we propose a new scheme for selecting a negative set of candidates for the retrieval task that can effectively improve the performance. We show that the proposed training strategy performs well on the task of change captioning on benchmark datasets.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b22bd1c6-2f4f-407c-836e-0e9cd8d0dcc1Cited by top-tier papers15
- Fine-tuning Multimodal LLMs to Follow Zero-shot Demonstrative InstructionsJuncheng Li, Kaihang Pan, Zhiqi Ge, Minghe Gao et al.ICLR 2024 · 95 citations
- Image Difference Captioning with Pre-training and Contrastive LearningLinli Yao, Weiying Wang, Qin JinAAAI 2022 · 66 citations
- Self-supervised Cross-view Representation Reconstruction for Change CaptioningYunbin Tu, Liang Li, Li Su, Zheng-Jun Zha et al.ICCV 2023 · 45 citations
- Image Retrieval from Contextual DescriptionsBenno Krojer, Vaibhav Adlakha, Vibhav Vineet, Yash Goyal et al.ACL 2022 · 37 citations
- Auxiliary Tasks Benefit 3D Skeleton-based Human Motion PredictionChenxin Xu, Robby T. Tan, Yuhong Tan, Siheng Chen et al.ICCV 2023 · 35 citations
Builds on9
- Attention on Attention for Image CaptioningLun Huang, Wenmin Wang, Jie Chen, Xiaoyong WeiICCV 2019 · 992 citations
- Entangled Transformer for Image CaptioningGuang Li, Linchao Zhu, Ping Liu, Yi YangICCV 2019 · 346 citations
- Robust Change CaptioningDong Huk Park, Trevor Darrell, Anna RohrbachICCV 2019 · 217 citations
- RATT: Recurrent Attention to Transient Tasks for Continual Image CaptioningRiccardo Del Chiaro, Bartlomiej Twardowski, Andrew D. Bagdanov, Joost van de WeijerNeurIPS 2020 · 55 citations
- Auxiliary Task Reweighting for Minimum-data LearningBaifeng Shi, Judy Hoffman, Kate Saenko, Trevor Darrell et al.NeurIPS 2020 · 44 citations
Related papers
- Context-aware Difference Distilling for Multi-change CaptioningYunbin Tu, Liang Li, Li Su, Zheng-Jun Zha et al.ACL 2024
- Retrieval Guided Unsupervised Multi-domain Image to Image TranslationRaul Gomez, Yahui Liu, Marco De Nadai, Dimosthenis Karatzas et al.ACM MM 2020 · 7 citations
- Differential-Perceptive and Retrieval-Augmented MLLM for Change CaptioningXian Zhang, Haokun Wen, Jianlong Wu, Pengda Qin et al.ACM MM 2024 · 6 citations
- Visual Delta Generator with Large Multi-Modal Models for Semi-Supervised Composed Image RetrievalYoung Kyun Jang, Donghyun Kim, Zihang Meng, Dat Huynh et al.CVPR 2024
- Cross-modal Joint Prediction and Alignment for Composed Query Image RetrievalYuchen Yang, Min Wang, Wengang Zhou, Houqiang LiACM MM 2021 · 29 citations
