Change3D: Revisiting Change Detection and Captioning from A Video Modeling Perspective
Duowang Zhu, Xiaohu Huang, Haiyan Huang, Hao Zhou, Zhenfeng Shao
摘要
In this paper, we present Change3D, a framework that reconceptualizes the change detection and captioning tasks through video modeling. Recent methods have achieved remarkable success by regarding each pair of bi-temporal images as separate frames. They employ a shared-weight image encoder to extract spatial features and then use a change extractor to capture differences between the two images. However, image feature encoding, being a taskagnostic process, cannot attend to changed regions effectively. Furthermore, different change extractors designed for various change detection and captioning tasks make it difficult to have a unified framework. To tackle these challenges, Change3D regards the bi-temporal images as comprising two frames akin to a tiny video. By integrating learnable perception frames between the bi-temporal images, a video encoder enables the perception frames to interact with the images directly and perceive their differences. Therefore, we can get rid of the intricate change extractors, providing a unified framework for different change detection and captioning tasks. We verify Change3D on multiple tasks, encompassing change detection (including binary change detection, semantic change detection, and building damage assessment) and change captioning, across eight standard benchmarks. Without bells and whistles, this simple yet effective framework can achieve superior performance with an ultra-light video model comprising only ∼6%-13% of the parameters and ∼8%-34% of the FLOPs compared to state-of-the-art methods. We hope that Change3D could be an alternative to 2D-based models and facilitate future research.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Sparsely Timing the Change: A Spiking Temporal Framework for Remote Sensing InterpretationShilong Li, Xiurui Xie, Qiugang Zhan, Luochao Wang 等CVPR 2026
- Exploring Generalizable Remote Sensing Change Detection via Low-Rank Exchange Adaptation of Vision Foundation ModelMingwei Zhang, Jingtao Hu, Qiang Li, Qi WangAAAI 2026
它引用的顶会 Paper19
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu 等ICCV 2021 · 被引用 31,683 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Pyramid Vision Transformer: A Versatile Backbone for Dense Prediction without ConvolutionsWenhai Wang, Enze Xie, Xiang Li, Deng-Ping Fan 等ICCV 2021 · 被引用 4,909 次
- SlowFast Networks for Video RecognitionChristoph Feichtenhofer, Haoqi Fan, Jitendra Malik, Kaiming HeICCV 2019 · 被引用 4,104 次
- Is Space-Time Attention All You Need for Video Understanding?Gedas Bertasius, Heng Wang, Lorenzo TorresaniICML 2021 · 被引用 2,927 次
相关 Paper
- Scene Graph with 3D Information for Change CaptioningZeming Liao, Qingbao Huang, Yu Liang, Mingyi Fu 等ACM MM 2021 · 被引用 17 次
- Describing and Localizing Multiple Changes with TransformersYue Qiu, Shintaro Yamamoto, Kodai Nakashima, Ryota Suzuki 等ICCV 2021 · 被引用 108 次
- R3Net: Relation-embedded Representation Reconstruction Network for Change CaptioningYunbin Tu, Liang Li, Chenggang Yan, Shengxiang Gao 等EMNLP 2021 · 被引用 21 次
- Imagine How To Change: Explicit Procedure Modeling for Change CaptioningJiayang Sun, Zixin Guo, Min Cao, Guibo Zhu 等ICLR 2026 · 被引用 1 次
- Viewpoint-Agnostic Change Captioning with Cycle ConsistencyHoeseong Kim, Jongseok Kim, Hyungseok Lee, Hyunsung Park 等ICCV 2021 · 被引用 56 次
