Describing and Localizing Multiple Changes with Transformers
Yue Qiu, Shintaro Yamamoto, Kodai Nakashima, Ryota Suzuki, Kenji Iwata, Hirokatsu Kataoka, Yutaka Satoh
Abstract
Change captioning tasks aim to detect changes in image pairs observed before and after a scene change and generate a natural language description of the changes. Existing change captioning studies have mainly focused on a single change. However, detecting and describing multiple changed parts in image pairs is essential for enhancing adaptability to complex scenarios. We solve the above issues from three aspects: (i) We propose a simulation-based multi-change captioning dataset; (ii) We benchmark existing state-of-the-art methods of single change captioning on multi-change captioning; (iii) We further propose Multi-Change Captioning transformers (MCCFormers) that identify change regions by densely correlating different regions in image pairs and dynamically determines the related change regions with words in sentences. The proposed method obtained the highest scores on four conventional change captioning evaluation metrics for multi-change captioning. Additionally, our proposed method can separate attention maps for each change and performs well with respect to change localization. Moreover, the proposed framework outperformed the previous state-of-the-art methods on an existing change captioning benchmark, CLEVR-Change, by a large margin (+6.1 on BLEU-4 and +9.7 on CIDEr scores), indicating its general ability in change captioning tasks. The code and dataset are available at the project page 1.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext f25b3229-bdf7-4572-97f4-c5c2d8a16fbfCited by top-tier papers14
- Self-supervised Cross-view Representation Reconstruction for Change CaptioningYunbin Tu, Liang Li, Li Su, Zheng-Jun Zha et al.ICCV 2023 · 45 citations
- Exploring Group Video Captioning with Efficient Relational ApproximationWang Lin, Tao Jin, Ye Wang, Wenwen Pan et al.ICCV 2023 · 17 citations
- Step Differences in Instructional VideoTushar Nagarajan, Lorenzo TorresaniCVPR 2024 · 7 citations
- The STVchrono Dataset: Towards Continuous Change Recognition in TimeYanjun Sun, Yue Qiu, Mariia Khan, Fumiya Matsuzawa et al.CVPR 2024 · 6 citations
- Region-aware Difference Distilling with Attribute-guided Contrastive Regularization for Change CaptioningRong Li, Liang Li, Jiehua Zhang, Qiang Zhao et al.AAAI 2025 · 4 citations
Builds on4
- Entangled Transformer for Image CaptioningGuang Li, Linchao Zhu, Ping Liu, Yi YangICCV 2019 · 346 citations
- Robust Change CaptioningDong Huk Park, Trevor Darrell, Anna RohrbachICCV 2019 · 217 citations
- Rescan: Inductive Instance Segmentation for Indoor RGBD ScansMaciej Halber, Yifei Shi, Kai Xu, Thomas A. FunkhouserICCV 2019 · 20 citations
- Meshed-Memory Transformer for Image CaptioningMarcella Cornia, Matteo Stefanini, Lorenzo Baraldi, Rita CucchiaraCVPR 2020
Related papers
- Context-aware Difference Distilling for Multi-change CaptioningYunbin Tu, Liang Li, Li Su, Zheng-Jun Zha et al.ACL 2024
- OmniDiff: A Comprehensive Benchmark for Fine-Grained Image Difference CaptioningYuan Liu, Saihui Hou, Saijie Hou, Jiabao Du et al.ICCV 2025 · 1 citation
- Scene Graph with 3D Information for Change CaptioningZeming Liao, Qingbao Huang, Yu Liang, Mingyi Fu et al.ACM MM 2021 · 17 citations
- Differential-Perceptive and Retrieval-Augmented MLLM for Change CaptioningXian Zhang, Haokun Wen, Jianlong Wu, Pengda Qin et al.ACM MM 2024 · 6 citations
- Viewpoint-Agnostic Change Captioning with Cycle ConsistencyHoeseong Kim, Jongseok Kim, Hyungseok Lee, Hyunsung Park et al.ICCV 2021 · 56 citations
