Image Change Captioning by Learning From an Auxiliary Task
Mehrdad Hosseinzadeh, Yang Wang
摘要
We tackle the challenging task of image change captioning. The goal is to describe the subtle difference between two very similar images by generating a sentence caption. While the recent methods mainly focus on proposing new model architectures for this problem, we instead focus on an alternative training scheme. Inspired by the success of multi-task learning, we formulate a training scheme that uses an auxiliary task to improve the training of the change captioning network. We argue that the task of composed query image retrieval is a natural choice as the auxiliary task. Given two almost similar images as the input, the primary network generates a caption describing the fine change between those two images. Next, the auxiliary network is provided with the generated caption and one of those two images. It then tries to pick the second image among a set of candidates. This forces the primary network to generate detailed and precise captions via having an extra supervision loss by the auxiliary network. Furthermore, we propose a new scheme for selecting a negative set of candidates for the retrieval task that can effectively improve the performance. We show that the proposed training strategy performs well on the task of change captioning on benchmark datasets.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper15
- Fine-tuning Multimodal LLMs to Follow Zero-shot Demonstrative InstructionsJuncheng Li, Kaihang Pan, Zhiqi Ge, Minghe Gao 等ICLR 2024 · 被引用 95 次
- Image Difference Captioning with Pre-training and Contrastive LearningLinli Yao, Weiying Wang, Qin JinAAAI 2022 · 被引用 66 次
- Self-supervised Cross-view Representation Reconstruction for Change CaptioningYunbin Tu, Liang Li, Li Su, Zheng-Jun Zha 等ICCV 2023 · 被引用 45 次
- Image Retrieval from Contextual DescriptionsBenno Krojer, Vaibhav Adlakha, Vibhav Vineet, Yash Goyal 等ACL 2022 · 被引用 37 次
- Auxiliary Tasks Benefit 3D Skeleton-based Human Motion PredictionChenxin Xu, Robby T. Tan, Yuhong Tan, Siheng Chen 等ICCV 2023 · 被引用 35 次
它引用的顶会 Paper9
- Attention on Attention for Image CaptioningLun Huang, Wenmin Wang, Jie Chen, Xiaoyong WeiICCV 2019 · 被引用 992 次
- Entangled Transformer for Image CaptioningGuang Li, Linchao Zhu, Ping Liu, Yi YangICCV 2019 · 被引用 346 次
- Robust Change CaptioningDong Huk Park, Trevor Darrell, Anna RohrbachICCV 2019 · 被引用 217 次
- RATT: Recurrent Attention to Transient Tasks for Continual Image CaptioningRiccardo Del Chiaro, Bartlomiej Twardowski, Andrew D. Bagdanov, Joost van de WeijerNeurIPS 2020 · 被引用 55 次
- Auxiliary Task Reweighting for Minimum-data LearningBaifeng Shi, Judy Hoffman, Kate Saenko, Trevor Darrell 等NeurIPS 2020 · 被引用 44 次
相关 Paper
- Context-aware Difference Distilling for Multi-change CaptioningYunbin Tu, Liang Li, Li Su, Zheng-Jun Zha 等ACL 2024
- Retrieval Guided Unsupervised Multi-domain Image to Image TranslationRaul Gomez, Yahui Liu, Marco De Nadai, Dimosthenis Karatzas 等ACM MM 2020 · 被引用 7 次
- Differential-Perceptive and Retrieval-Augmented MLLM for Change CaptioningXian Zhang, Haokun Wen, Jianlong Wu, Pengda Qin 等ACM MM 2024 · 被引用 6 次
- Visual Delta Generator with Large Multi-Modal Models for Semi-Supervised Composed Image RetrievalYoung Kyun Jang, Donghyun Kim, Zihang Meng, Dat Huynh 等CVPR 2024
- Cross-modal Joint Prediction and Alignment for Composed Query Image RetrievalYuchen Yang, Min Wang, Wengang Zhou, Houqiang LiACM MM 2021 · 被引用 29 次
