VSCD: Video-based Scene Change Detection in Unaligned Scenes
Jiae Yoon, Ue-Hwan Kim
Abstract
Detecting what has changed in an environment is essential for long-term autonomy, yet most change detection settings assume fixed viewpoints, mild misalignment, or only a few changed objects. We introduce Video-based Scene Change Detection (VSCD), which predicts a pixel-wise change mask for each query frame, given a reference and a query RGB video of the same indoor space recorded at different times under unconstrained camera motion. The two videos are not temporally synchronized, and many object instances may appear or disappear. To study this setting, we build a large-scale benchmark with over 1.1 million frames annotated with pixel-accurate change masks, together with a real-world test set for evaluating transfer beyond simulation. We propose a query-centric multi-reference model that learns temporal matching implicitly from change-mask supervision, aligns candidate reference features to the query via local patch correspondence, and fuses per-candidate change features using frame-level and patch-level confidence before decoding a high-resolution mask once per frame. Our approach achieves state-of-the-art performance against strong image- and video-based baselines, and we validate its real-world impact by deploying it on a mobile robot for two downstream applications—visual surveillance and object incremental learning.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 8da7186d-c628-4eed-b530-cba1589f0ee6Builds on6
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao et al.ICCV 2023 · 13,211 citations
- PaLM-E: An Embodied Multimodal Language ModelDanny Driess, Fei Xia, Mehdi S. M. Sajjadi, Corey Lynch et al.ICML 2023 · 2,601 citations
- Video Similarity and Alignment Learning on Partial Video Copy DetectionZhen Han, Xiangteng He, Mingqian Tang, Yiliang LvACM MM 2021 · 34 citations
- TransVCL: Attention-Enhanced Video Copy Localization Network with Flexible SupervisionSifeng He, Yue He, Minlong Lu, Chen Jiang et al.AAAI 2023 · 26 citations
- Zero-Shot Scene Change DetectionKyusik Cho, Dong Yeop Kim, Euntai KimAAAI 2025 · 10 citations
Related papers
- Dual Task Learning by Leveraging Both Dense Correspondence and Mis-Correspondence for Robust Change Detection With Imperfect MatchesJin-Man Park, Ue-Hwan Kim, Seon-Hoon Lee, Jong-Hwan KimCVPR 2022 · 13 citations
- Towards Generalizable Scene Change DetectionJae-Woo Kim, Ue-Hwan KimCVPR 2025
- Changes in Real Time: Online Scene Change Detection with Multi-View FusionChamuditha Jayanga Galappaththige, Jason Lai, Lloyd Windrim, Donald G. Dansereau et al.CVPR 2026 · 4 citations
- 3D Video Object Detection with Learnable Object-Centric Global OptimizationJiawei He, Yuntao Chen, Naiyan Wang, Zhaoxiang ZhangCVPR 2023
- Multi-View Pose-Agnostic Change Localization with Zero LabelsChamuditha Jayanga Galappaththige, Jason Lai, Lloyd Windrim, Donald G. Dansereau et al.CVPR 2025
