Self-supervised Cross-view Representation Reconstruction for Change Captioning
Yunbin Tu, Liang Li, Li Su, Zheng-Jun Zha, Chenggang Yan, Qingming Huang
Abstract
Change captioning aims to describe the difference between a pair of similar images. Its key challenge is how to learn a stable difference representation under pseudo changes caused by viewpoint change. In this paper, we address this by proposing a self-supervised cross-view representation reconstruction (SCORER) network. Concretely, we first design a multi-head token-wise matching to model relationships between cross-view features from similar/dissimilar images. Then, by maximizing cross-view contrastive alignment of two similar images, SCORER learns two view-invariant image representations in a self-supervised way. Based on these, we reconstruct the representations of unchanged objects by cross-attention, thus learning a stable difference representation for caption generation. Further, we devise a cross-modal backward reasoning to improve the quality of caption. This module reversely models a "hallucination" representation with the caption and "before" representation. By pushing it closer to the "after" representation, we enforce the caption to be informative about the difference in a self-supervised manner. Extensive experiments show our method achieves the state-of-the-art results on four datasets. The code is available at https://github.com/tuyunbin/SCORER.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 0ef86c75-5fa0-4b5a-8e09-3e05164ad9f3Cited by top-tier papers17
- R&B: Region and Boundary Aware Zero-shot Grounded Text-to-image GenerationJiayu Xiao, Henglei Lv, Liang Li, Shuhui Wang et al.ICLR 2024 · 31 citations
- Synergistic Multi-Agent Framework with Trajectory Learning for Knowledge-Intensive TasksShengbin Yue, Siyuan Wang, Wei Chen, Xuanjing Huang et al.AAAI 2025 · 25 citations
- The STVchrono Dataset: Towards Continuous Change Recognition in TimeYanjun Sun, Yue Qiu, Mariia Khan, Fumiya Matsuzawa et al.CVPR 2024 · 6 citations
- Exploring Temporal Event Cues for Dense Video Captioning in Cyclic Co-LearningZhuyang Xie, Yan Yang, Yankai Yu, Jie Wang et al.AAAI 2025 · 5 citations
- Region-aware Difference Distilling with Attribute-guided Contrastive Regularization for Change CaptioningRong Li, Liang Li, Jiehua Zhang, Qiang Zhao et al.AAAI 2025 · 4 citations
Builds on10
- FILIP: Fine-grained Interactive Language-Image Pre-TrainingLewei Yao, Runhui Huang, Lu Hou, Guansong Lu et al.ICLR 2022 · 827 citations
- SwinBERT: End-to-End Transformers with Sparse Attention for Video CaptioningKevin Lin, Linjie Li, Chung-Ching Lin, Faisal Ahmed et al.CVPR 2022 · 263 citations
- Robust Change CaptioningDong Huk Park, Trevor Darrell, Anna RohrbachICCV 2019 · 217 citations
- Describing and Localizing Multiple Changes with TransformersYue Qiu, Shintaro Yamamoto, Kodai Nakashima, Ryota Suzuki et al.ICCV 2021 · 108 citations
- Image Difference Captioning with Pre-training and Contrastive LearningLinli Yao, Weiying Wang, Qin JinAAAI 2022 · 66 citations
Related papers
- Revisiting Change Captioning from Self-supervised Global-Part AlignmentFeixiao Lv, Rui Wang, Lihua JingAAAI 2025 · 1 citation
- R3Net: Relation-embedded Representation Reconstruction Network for Change CaptioningYunbin Tu, Liang Li, Chenggang Yan, Shengxiang Gao et al.EMNLP 2021 · 21 citations
- Scene Graph with 3D Information for Change CaptioningZeming Liao, Qingbao Huang, Yu Liang, Mingyi Fu et al.ACM MM 2021 · 17 citations
- DECIDER: Difference-aware Contrastive Diffusion Model with Adversarial Perturbations for Image Change CaptioningGuojin Zhong, Jinhong Hu, Jiajun Chen, Jin Yuan et al.AAAI 2025 · 3 citations
- Viewpoint-Agnostic Change Captioning with Cycle ConsistencyHoeseong Kim, Jongseok Kim, Hyungseok Lee, Hyunsung Park et al.ICCV 2021 · 56 citations
