Image Difference Captioning via Adversarial Preference Optimization
Zihan Huang, Junda Wu, Rohan Surana, Tong Yu, David Arbour, Ritwik Sinha, Julian J. McAuley
摘要
Image Difference Captioning (IDC) aims to generate natural language descriptions that highlight subtle differences between two visually similar images. While recent advances leverage pre-trained vision-language models to align fine-grained visual differences with textual semantics, existing supervised approaches often overly focus on dataset-specific language patterns and fail to capture fine-grained and context-aware preferences on IDC, due to limited annotation diversity and a lack of semantically informative negative examples during training, To address these limitations, we propose an adversarial direct preference optimization (ADPO) framework for IDC, which formulates IDC as a preference optimization problem under the Bradley-Terry-Luce model, directly aligning the captioning policy with pairwise difference preferences via Direct Preference Optimization (DPO). To model more accurate and diverse IDC preferences, we introduce an adversarially trained hard negative retriever that selects counterfactual captions, This results in a minimax optimization problem, which we solve via policy-gradient reinforcement learning, enabling the policy and retriever to improve jointly. By dynamically generating semantically challenging negatives, our method reduces reliance on dataset-specific patterns. Experiments on benchmark IDC datasets show that our approach outperforms existing baselines, especially in generating fine-grained and accurate difference descriptions. * These authors contributed equally. Query: Please describe what the difference is between the target image and the reference image Reference Image Target Image Chosen: "the person is folding a green paper in right image" Rejected: "the blue truck is now in the picture on the right" GT IDC:"the person is folding a green paper in right image" SFT can overly focus on dataset-specific language patterns, e.g., "the person", "right image". DPO with trivial comparisons fail to learn subtle differences Chosen: "the person is folding a green paper in right image" Rejected: "the person is folding a red paper in right image" Adversarial DPO with adversarial learned negative retrieval benefits to learn more difficult IDC with nuanced difference
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper13
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 被引用 11,349 次
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning 等NeurIPS 2023 · 被引用 10,924 次
- Robust Change CaptioningDong Huk Park, Trevor Darrell, Anna RohrbachICCV 2019 · 被引用 217 次
- Learning Efficient Vision Transformers via Fine-Grained Manifold DistillationZhiwei Hao, Jianyuan Guo, Ding Jia, Kai Han 等NeurIPS 2022 · 被引用 103 次
- Image Difference Captioning with Pre-training and Contrastive LearningLinli Yao, Weiying Wang, Qin JinAAAI 2022 · 被引用 66 次
相关 Paper
- DECIDER: Difference-aware Contrastive Diffusion Model with Adversarial Perturbations for Image Change CaptioningGuojin Zhong, Jinhong Hu, Jiajun Chen, Jin Yuan 等AAAI 2025 · 被引用 3 次
- Region-aware Difference Distilling with Attribute-guided Contrastive Regularization for Change CaptioningRong Li, Liang Li, Jiehua Zhang, Qiang Zhao 等AAAI 2025 · 被引用 4 次
- Diffusion Model Alignment Using Direct Preference OptimizationBram Wallace, Meihua Dang, Rafael Rafailov, Linqi Zhou 等CVPR 2024 · 被引用 89 次
- Towards Fine-Grained Attribution: Instance-Aware Preference Optimization for Aligning Diffusion ModelsJiayang Sun, Pin Wang, Hongbo Wang, Xinyue Liu 等CVPR 2026
- Ranking-based Preference Optimization for Diffusion Models from Implicit User FeedbackYi-Lun Wu, Bo-Kai Ruan, Chiang Tseng, Hong-Han ShuaiNeurIPS 2025 · 被引用 3 次
