Differential-Perceptive and Retrieval-Augmented MLLM for Change Captioning
Xian Zhang, Haokun Wen, Jianlong Wu, Pengda Qin, Hui Xue', Liqiang Nie
摘要
Change captioning involves describing the subtle changes between a pair of similar images. Although existing efforts have achieved compelling success, they overlook the potential of multimodal large language models (MLLMs) in tackling this challenging task. In this work, we aim to empower MLLMs with the capability to perceive subtle differences between paired images and enhance their performance in generating change captions. Specifically, we present a diFferentIal-perceptive aNd rEtRieval-augmented MLLM (FINER-MLLM) tailored for this task. In particular, FINER-MLLM leverages LoRA fine-tuned MLLM's image encoder to extract image patch features, enabling the capture of detailed image information. Subsequently, within MLLM's feature extraction, typically Q-Former, FINER-MLLM incorporates dual constraints: the intra-image feature independence constraint and the inter-image feature alignment constraint. These constraints ensure that the features can comprehensively extract subtle visual information within each image and that corresponding features across images align effectively. Last, we introduced the retrieval augmentation to first retrieve the relevant corpus to facilitate the MLLM's decoder i.e., LLM, in generating accurate change captions. Extensive experiments on three benchmark datasets, i.e., CLEVR-Change, Spot-the-Diff, and Image-Editing-Request, demonstrate the superiority of our proposed method.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
引用它的顶会 Paper5
- MIDAS: Multi-Image Dispersion and Semantic Reconstruction for Jailbreaking MLLMsYilian Liu, Guoshun Nan, Jiuyang Lyu, Zhican Chen 等ICLR 2026 · 被引用 3 次
- Image Difference Captioning via Adversarial Preference OptimizationZihan Huang, Junda Wu, Rohan Surana, Tong Yu 等EMNLP 2025 · 被引用 3 次
- OmniDiff: A Comprehensive Benchmark for Fine-Grained Image Difference CaptioningYuan Liu, Saihui Hou, Saijie Hou, Jiabao Du 等ICCV 2025 · 被引用 1 次
- Imagine How To Change: Explicit Procedure Modeling for Change CaptioningJiayang Sun, Zixin Guo, Min Cao, Guibo Zhu 等ICLR 2026 · 被引用 1 次
- MP-GUI: Modality Perception with MLLMs for GUI UnderstandingZiwei Wang, Weizhi Chen, Leyang Yang, Sheng Zhou 等CVPR 2025
相关 Paper
- Describing and Localizing Multiple Changes with TransformersYue Qiu, Shintaro Yamamoto, Kodai Nakashima, Ryota Suzuki 等ICCV 2021 · 被引用 108 次
- Img-Diff: Contrastive Data Synthesis for Multimodal Large Language ModelsQirui Jiao, Daoyuan Chen, Yilun Huang, Bolin Ding 等CVPR 2025
- Hallucination at a Glance: Controlled Visual Edits and Fine-Grained Multimodal LearningTianyi Bai, Yuxuan Fan, Jiantao Qiu, Fupeng Sun 等NeurIPS 2025 · 被引用 12 次
- DECIDER: Difference-aware Contrastive Diffusion Model with Adversarial Perturbations for Image Change CaptioningGuojin Zhong, Jinhong Hu, Jiajun Chen, Jin Yuan 等AAAI 2025 · 被引用 3 次
- Image Change Captioning by Learning From an Auxiliary TaskMehrdad Hosseinzadeh, Yang WangCVPR 2021
