Viewpoint-Agnostic Change Captioning with Cycle Consistency
Hoeseong Kim, Jongseok Kim, Hyungseok Lee, Hyunsung Park, Gunhee Kim
Abstract
Change captioning is the task of identifying the change and describing it with a concise caption. Despite recent advancements, filtering out insignificant changes still remains as a challenge. Namely, images from different camera perspectives can cause issues; a mere change in viewpoint should be disregarded while still capturing the actual changes. In order to tackle this problem, we present a new Viewpoint-Agnostic change captioning network with Cycle Consistency (VACC) that requires only one image each for the before and after scene, without depending on any other information. We achieve this by devising a new difference encoder module which can encode viewpoint information and model the difference more effectively. In addition, we propose a cycle consistency module that can potentially improve the performance of any change captioning networks in general by matching the composite feature of the generated caption and before image with the after image feature. We evaluate the performance of our proposed model across three datasets for change captioning, including a novel dataset we introduce here that contains images with changes under extreme viewpoint shifts. Through our experiments, we show the excellence of our method with respect to the CIDEr, BLEU-4, METEOR and SPICE scores. Moreover, we demonstrate that attaching our proposed cycle consistency module yields a performance boost for existing change captioning networks, even with varying image encoding mechanisms.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 5b662617-c168-4499-afd5-92f2ec87bb80Cited by top-tier papers7
- Video Action DifferencingJames Burgess, Xiaohan Wang, Yuhui Zhang, Anita Rau et al.ICLR 2025 · 1,149 citations
- Exploring Temporal Event Cues for Dense Video Captioning in Cyclic Co-LearningZhuyang Xie, Yan Yang, Yankai Yu, Jie Wang et al.AAAI 2025 · 5 citations
- Image Difference Captioning via Adversarial Preference OptimizationZihan Huang, Junda Wu, Rohan Surana, Tong Yu et al.EMNLP 2025 · 3 citations
- DiffTell: A High-Quality Dataset for Describing Image Manipulation ChangesZonglin Di, Jing Shi, Yifei Fan, Hao Tan et al.ICCV 2025 · 1 citation
- Explicit Temporal-Semantic Modeling for Dense Video Captioning via Context-Aware Cross-Modal InteractionMingda Jia, Weiliang Meng, Zenghuang Fu, Yiheng Li et al.AAAI 2026 · 1 citation
Builds on3
- Robust Change CaptioningDong Huk Park, Trevor Darrell, Anna RohrbachICCV 2019 · 217 citations
- Dual Compositional Learning in Interactive Image RetrievalJongseok Kim, Youngjae Yu, Hoeseong Kim, Gunhee KimAAAI 2021 · 116 citations
- Few-Shot Generalization for Single-Image 3D Reconstruction via PriorsBram Wallace, Bharath HariharanICCV 2019 · 43 citations
Related papers
- Context-aware Difference Distilling for Multi-change CaptioningYunbin Tu, Liang Li, Li Su, Zheng-Jun Zha et al.ACL 2024
- Scene Graph with 3D Information for Change CaptioningZeming Liao, Qingbao Huang, Yu Liang, Mingyi Fu et al.ACM MM 2021 · 17 citations
- R3Net: Relation-embedded Representation Reconstruction Network for Change CaptioningYunbin Tu, Liang Li, Chenggang Yan, Shengxiang Gao et al.EMNLP 2021 · 21 citations
- Describing and Localizing Multiple Changes with TransformersYue Qiu, Shintaro Yamamoto, Kodai Nakashima, Ryota Suzuki et al.ICCV 2021 · 108 citations
- Self-supervised Cross-view Representation Reconstruction for Change CaptioningYunbin Tu, Liang Li, Li Su, Zheng-Jun Zha et al.ICCV 2023 · 45 citations
