R3Net: Relation-embedded Representation Reconstruction Network for Change Captioning
Yunbin Tu, Liang Li, Chenggang Yan, Shengxiang Gao, Zhengtao Yu
Abstract
Change captioning is to use a natural language sentence to describe the fine-grained disagreement between two similar images. Viewpoint change is the most typical distractor in this task, because it changes the scale and location of the objects and overwhelms the representation of real change. In this paper, we propose a Relation-embedded Representation Reconstruction Network (R 3 Net) to explicitly distinguish the real change from the large amount of clutter and irrelevant changes. Specifically, a relation-embedded module is first devised to explore potential changed objects in the large amount of clutter. Then, based on the semantic similarities of corresponding locations in the two images, a representation reconstruction module (RRM) is designed to learn the reconstruction representation and further model the difference representation. Besides, we introduce a syntactic skeleton predictor (SSP) to enhance the semantic interaction between change localization and caption generation. Extensive experiments show that the proposed method achieves the state-of-the-art results on two public datasets 1 .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext fec37271-c7cc-42ee-abfc-3cae71222fa1Cited by top-tier papers7
- Self-supervised Cross-view Representation Reconstruction for Change CaptioningYunbin Tu, Liang Li, Li Su, Zheng-Jun Zha et al.ICCV 2023 · 45 citations
- Improving Image Captioning via Predicting Structured ConceptsTing Wang, Weidong Chen, Yuanhe Tian, Yan Song et al.EMNLP 2023 · 16 citations
- Exploring Temporal Event Cues for Dense Video Captioning in Cyclic Co-LearningZhuyang Xie, Yan Yang, Yankai Yu, Jie Wang et al.AAAI 2025 · 5 citations
- Region-aware Difference Distilling with Attribute-guided Contrastive Regularization for Change CaptioningRong Li, Liang Li, Jiehua Zhang, Qiang Zhao et al.AAAI 2025 · 4 citations
- Image Difference Captioning via Adversarial Preference OptimizationZihan Huang, Junda Wu, Rohan Surana, Tong Yu et al.EMNLP 2025 · 3 citations
Builds on5
- Robust Change CaptioningDong Huk Park, Trevor Darrell, Anna RohrbachICCV 2019 · 217 citations
- A Novel Graph-based Multi-modal Fusion Encoder for Neural Machine TranslationYongjing Yin, Fandong Meng, Jinsong Su, Chulun Zhou et al.ACL 2020 · 145 citations
- Aligned Dual Channel Graph Convolutional Network for Visual Question AnsweringQingbao Huang, Jielong Wei, Yi Cai, Changmeng Zheng et al.ACL 2020 · 79 citations
- IR-GAN: Image Manipulation with Linguistic Instruction by Increment ReasoningZhenhuan Liu, Jincan Deng, Liang Li, Shaofei Cai et al.ACM MM 2020 · 17 citations
- Diverter-Guider Recurrent Network for Diverse Poems Generation from ImageLiang Li, Shijie Yang, Li Su, Shuhui Wang et al.ACM MM 2020 · 7 citations
Related papers
- Scene Graph with 3D Information for Change CaptioningZeming Liao, Qingbao Huang, Yu Liang, Mingyi Fu et al.ACM MM 2021 · 17 citations
- Context-aware Difference Distilling for Multi-change CaptioningYunbin Tu, Liang Li, Li Su, Zheng-Jun Zha et al.ACL 2024
- Revisiting Change Captioning from Self-supervised Global-Part AlignmentFeixiao Lv, Rui Wang, Lihua JingAAAI 2025 · 1 citation
- Viewpoint-Agnostic Change Captioning with Cycle ConsistencyHoeseong Kim, Jongseok Kim, Hyungseok Lee, Hyunsung Park et al.ICCV 2021 · 56 citations
- DECIDER: Difference-aware Contrastive Diffusion Model with Adversarial Perturbations for Image Change CaptioningGuojin Zhong, Jinhong Hu, Jiajun Chen, Jin Yuan et al.AAAI 2025 · 3 citations
