Differential-Perceptive and Retrieval-Augmented MLLM for Change Captioning
Xian Zhang, Haokun Wen, Jianlong Wu, Pengda Qin, Hui Xue', Liqiang Nie
Abstract
Change captioning involves describing the subtle changes between a pair of similar images. Although existing efforts have achieved compelling success, they overlook the potential of multimodal large language models (MLLMs) in tackling this challenging task. In this work, we aim to empower MLLMs with the capability to perceive subtle differences between paired images and enhance their performance in generating change captions. Specifically, we present a diFferentIal-perceptive aNd rEtRieval-augmented MLLM (FINER-MLLM) tailored for this task. In particular, FINER-MLLM leverages LoRA fine-tuned MLLM's image encoder to extract image patch features, enabling the capture of detailed image information. Subsequently, within MLLM's feature extraction, typically Q-Former, FINER-MLLM incorporates dual constraints: the intra-image feature independence constraint and the inter-image feature alignment constraint. These constraints ensure that the features can comprehensively extract subtle visual information within each image and that corresponding features across images align effectively. Last, we introduced the retrieval augmentation to first retrieve the relevant corpus to facilitate the MLLM's decoder i.e., LLM, in generating accurate change captions. Extensive experiments on three benchmark datasets, i.e., CLEVR-Change, Spot-the-Diff, and Image-Editing-Request, demonstrate the superiority of our proposed method.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get d9f4026d-d5c6-47e8-9992-ee1536c41342Cited by top-tier papers5
- MIDAS: Multi-Image Dispersion and Semantic Reconstruction for Jailbreaking MLLMsYilian Liu, Guoshun Nan, Jiuyang Lyu, Zhican Chen et al.ICLR 2026 · 3 citations
- Image Difference Captioning via Adversarial Preference OptimizationZihan Huang, Junda Wu, Rohan Surana, Tong Yu et al.EMNLP 2025 · 3 citations
- OmniDiff: A Comprehensive Benchmark for Fine-Grained Image Difference CaptioningYuan Liu, Saihui Hou, Saijie Hou, Jiabao Du et al.ICCV 2025 · 1 citation
- Imagine How To Change: Explicit Procedure Modeling for Change CaptioningJiayang Sun, Zixin Guo, Min Cao, Guibo Zhu et al.ICLR 2026 · 1 citation
- MP-GUI: Modality Perception with MLLMs for GUI UnderstandingZiwei Wang, Weizhi Chen, Leyang Yang, Sheng Zhou et al.CVPR 2025
Related papers
- Describing and Localizing Multiple Changes with TransformersYue Qiu, Shintaro Yamamoto, Kodai Nakashima, Ryota Suzuki et al.ICCV 2021 · 108 citations
- Img-Diff: Contrastive Data Synthesis for Multimodal Large Language ModelsQirui Jiao, Daoyuan Chen, Yilun Huang, Bolin Ding et al.CVPR 2025
- Hallucination at a Glance: Controlled Visual Edits and Fine-Grained Multimodal LearningTianyi Bai, Yuxuan Fan, Jiantao Qiu, Fupeng Sun et al.NeurIPS 2025 · 12 citations
- DECIDER: Difference-aware Contrastive Diffusion Model with Adversarial Perturbations for Image Change CaptioningGuojin Zhong, Jinhong Hu, Jiajun Chen, Jin Yuan et al.AAAI 2025 · 3 citations
- Image Change Captioning by Learning From an Auxiliary TaskMehrdad Hosseinzadeh, Yang WangCVPR 2021
