DECIDER: Difference-aware Contrastive Diffusion Model with Adversarial Perturbations for Image Change Captioning
Guojin Zhong, Jinhong Hu, Jiajun Chen, Jin Yuan, Wenbo Pan
摘要
Image change captioning (ICC) poses great challenges stemming from describing subtle differences between two similar images in natural language, significantly increasing the complexity of feature extraction and cross-modal learning compared to the image captioning task. Existing ICC methods often suffer from two key challenges: 1) Massive irrelevant information of uni-image features leads to suboptimal visual difference representations; 2) Imprecise inter-modality correspondence degrades the quality of generated captions. This paper proposes a Difference-aware Contrastive Diffusion Model with Adversarial Perturbations (DECIDER) for ICC due to the excellent performance of diffusion models in image/text generation. Technically, difference-aware cross-modal learning is developed to suppress irrelevant information and learn compact yet robust visual difference representations. This is achieved by optimizing a novel objective mathematically derived from the information bottleneck principle that excels in filtering redundant features and highlighting differences. Furthermore, we propose to dynamically generate ``hard'' positive and negative samples via adversarial perturbations, which are involved in contrastive diffusion training with a tighter variational bound. This design encourages our DECIDER to excavate and construct complex correspondences between visual differences and captions, thereby improving generalization performance. Extensive experiments on four datasets demonstrate that DECIDER significantly exceeds state-of-the-art performance.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Leveraging Textual Compositional Reasoning for Robust Change CaptioningKyu Ri Park, Jiyoung Park, Seong Tae Kim, Hong Joo Lee 等AAAI 2026
- Multi-Resolution Decomposable Diffusion Model for Non-Stationary Time Series Anomaly DetectionGuojin Zhong, Pan Wang, Jin Yuan, Zhiyong Li 等ICLR 2025
它引用的顶会 Paper18
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 被引用 11,349 次
- Structured Denoising Diffusion Models in Discrete State-SpacesJacob Austin, Daniel D. Johnson, Jonathan Ho, Daniel Tarlow 等NeurIPS 2021 · 被引用 2,256 次
- Vector Quantized Diffusion Model for Text-to-Image SynthesisShuyang Gu, Dong Chen, Jianmin Bao, Fang Wen 等CVPR 2022 · 被引用 607 次
- CLUB: A Contrastive Log-ratio Upper Bound of Mutual InformationPengyu Cheng, Weituo Hao, Shuyang Dai, Jiachang Liu 等ICML 2020 · 被引用 512 次
- Robust Change CaptioningDong Huk Park, Trevor Darrell, Anna RohrbachICCV 2019 · 被引用 217 次
相关 Paper
- Noise-Aware Image Captioning with Progressively Exploring Mismatched WordsZhongtian Fu, Kefei Song, Luping Zhou, Yang YangAAAI 2024 · 被引用 36 次
- Self-supervised Cross-view Representation Reconstruction for Change CaptioningYunbin Tu, Liang Li, Li Su, Zheng-Jun Zha 等ICCV 2023 · 被引用 45 次
- Region-aware Difference Distilling with Attribute-guided Contrastive Regularization for Change CaptioningRong Li, Liang Li, Jiehua Zhang, Qiang Zhao 等AAAI 2025 · 被引用 4 次
- Semantic-Conditional Diffusion Networks for Image CaptioningJianjie Luo, Yehao Li, Yingwei Pan, Ting Yao 等CVPR 2023
- Contrast-augmented Diffusion Model with Fine-grained Sequence Alignment for Markup-to-Image GenerationGuojin Zhong, Jin Yuan, Pan Wang, Kailun Yang 等ACM MM 2023 · 被引用 7 次
