Self-supervised Cross-view Representation Reconstruction for Change Captioning
Yunbin Tu, Liang Li, Li Su, Zheng-Jun Zha, Chenggang Yan, Qingming Huang
摘要
Change captioning aims to describe the difference between a pair of similar images. Its key challenge is how to learn a stable difference representation under pseudo changes caused by viewpoint change. In this paper, we address this by proposing a self-supervised cross-view representation reconstruction (SCORER) network. Concretely, we first design a multi-head token-wise matching to model relationships between cross-view features from similar/dissimilar images. Then, by maximizing cross-view contrastive alignment of two similar images, SCORER learns two view-invariant image representations in a self-supervised way. Based on these, we reconstruct the representations of unchanged objects by cross-attention, thus learning a stable difference representation for caption generation. Further, we devise a cross-modal backward reasoning to improve the quality of caption. This module reversely models a "hallucination" representation with the caption and "before" representation. By pushing it closer to the "after" representation, we enforce the caption to be informative about the difference in a self-supervised manner. Extensive experiments show our method achieves the state-of-the-art results on four datasets. The code is available at https://github.com/tuyunbin/SCORER.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper17
- R&B: Region and Boundary Aware Zero-shot Grounded Text-to-image GenerationJiayu Xiao, Henglei Lv, Liang Li, Shuhui Wang 等ICLR 2024 · 被引用 31 次
- Synergistic Multi-Agent Framework with Trajectory Learning for Knowledge-Intensive TasksShengbin Yue, Siyuan Wang, Wei Chen, Xuanjing Huang 等AAAI 2025 · 被引用 25 次
- The STVchrono Dataset: Towards Continuous Change Recognition in TimeYanjun Sun, Yue Qiu, Mariia Khan, Fumiya Matsuzawa 等CVPR 2024 · 被引用 6 次
- Exploring Temporal Event Cues for Dense Video Captioning in Cyclic Co-LearningZhuyang Xie, Yan Yang, Yankai Yu, Jie Wang 等AAAI 2025 · 被引用 5 次
- Region-aware Difference Distilling with Attribute-guided Contrastive Regularization for Change CaptioningRong Li, Liang Li, Jiehua Zhang, Qiang Zhao 等AAAI 2025 · 被引用 4 次
它引用的顶会 Paper10
- FILIP: Fine-grained Interactive Language-Image Pre-TrainingLewei Yao, Runhui Huang, Lu Hou, Guansong Lu 等ICLR 2022 · 被引用 827 次
- SwinBERT: End-to-End Transformers with Sparse Attention for Video CaptioningKevin Lin, Linjie Li, Chung-Ching Lin, Faisal Ahmed 等CVPR 2022 · 被引用 263 次
- Robust Change CaptioningDong Huk Park, Trevor Darrell, Anna RohrbachICCV 2019 · 被引用 217 次
- Describing and Localizing Multiple Changes with TransformersYue Qiu, Shintaro Yamamoto, Kodai Nakashima, Ryota Suzuki 等ICCV 2021 · 被引用 108 次
- Image Difference Captioning with Pre-training and Contrastive LearningLinli Yao, Weiying Wang, Qin JinAAAI 2022 · 被引用 66 次
相关 Paper
- Revisiting Change Captioning from Self-supervised Global-Part AlignmentFeixiao Lv, Rui Wang, Lihua JingAAAI 2025 · 被引用 1 次
- R3Net: Relation-embedded Representation Reconstruction Network for Change CaptioningYunbin Tu, Liang Li, Chenggang Yan, Shengxiang Gao 等EMNLP 2021 · 被引用 21 次
- Scene Graph with 3D Information for Change CaptioningZeming Liao, Qingbao Huang, Yu Liang, Mingyi Fu 等ACM MM 2021 · 被引用 17 次
- DECIDER: Difference-aware Contrastive Diffusion Model with Adversarial Perturbations for Image Change CaptioningGuojin Zhong, Jinhong Hu, Jiajun Chen, Jin Yuan 等AAAI 2025 · 被引用 3 次
- Viewpoint-Agnostic Change Captioning with Cycle ConsistencyHoeseong Kim, Jongseok Kim, Hyungseok Lee, Hyunsung Park 等ICCV 2021 · 被引用 56 次
