The STVchrono Dataset: Towards Continuous Change Recognition in Time
Yanjun Sun, Yue Qiu, Mariia Khan, Fumiya Matsuzawa, Kenji Iwata
Abstract
Recognizing continuous changes offers valuable insights into past historical events, supports current trend analysis, and facilitates future planning. This knowledge is crucial for a variety of fields, such as meteorology and agriculture, environmental science, urban planning and construction, tourism, and cultural preservation. Currently available datasets in the field of scene change understanding primarily concentrate on two main tasks: the detection of changed regions within a scene and the linguistic description of the change content. Existing datasets focus on recognizing discrete changes, such as adding or deleting an object from two images, and largely rely on artificially generated images. Consequently, the existing change understanding methods primarily focus on identifying distinct object differences, overlooking the importance of continuous, gradual changes occurring over extended time intervals. To address the above issues, we propose a novel benchmark dataset, STVchrono, targeting the localization and description of long-term continuous changes in real-world scenes. The dataset consists of 71,900 photographs from Google Street View API taken over an 18-year span across 50 cities all over the world. Our STVchrono dataset is designed to support real-world continuous change recognition and description in both image pairs and extended image sequences, while also enabling the segmentation of changed regions. We conduct experiments to evaluate state-of-the- art methods on continuous change description and segmentation, as well as multimodal Large Language Models for describing changes. Our findings reveal that even the most advanced methods lag human performance, emphasizing the need to adapt them to continuously changing real-world scenarios. We hope that our benchmark dataset will further facilitate the research of temporal change recognition in a dynamic world. The STVchrono dataset is available at STVchrono Dataset.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext ff05c498-01c5-4dd8-b325-dcb3224c3e65Cited by top-tier papers5
- TAB: Transformer Attention Bottlenecks Enable User Intervention and Debugging in Vision-Language ModelsPooyan Rahmanzadehgervi, Hung Huy Nguyen, Rosanne Liu, Long Mai et al.ICCV 2025 · 3 citations
- OmniDiff: A Comprehensive Benchmark for Fine-Grained Image Difference CaptioningYuan Liu, Saihui Hou, Saijie Hou, Jiabao Du et al.ICCV 2025 · 1 citation
- Imagine How To Change: Explicit Procedure Modeling for Change CaptioningJiayang Sun, Zixin Guo, Min Cao, Guibo Zhu et al.ICLR 2026 · 1 citation
- VideoSetDiff: Identifying and Reasoning Similarities and Differences in Similar VideosYue Qiu, Yanjun Sun, Takuma Yagi, Shusaku Egami et al.ICCV 2025 · 1 citation
- STATUS Bench: A Rigorous Benchmark for Evaluating Object State Understanding in Vision-Language ModelsMahiro Ukai, Shuhei Kurita, Nakamasa InoueACM MM 2025
Builds on19
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 7,873 citations
- Per-Pixel Classification is Not All You Need for Semantic SegmentationBowen Cheng, Alexander G. Schwing, Alexander KirillovNeurIPS 2021 · 2,196 citations
- Robust Change CaptioningDong Huk Park, Trevor Darrell, Anna RohrbachICCV 2019 · 217 citations
- SatlasPretrain: A Large-Scale Dataset for Remote Sensing Image UnderstandingFavyen Bastani, Piper Wolters, Ritwik Gupta, Joe Ferdinando et al.ICCV 2023 · 216 citations
Related papers
- Mapillary Street-Level Sequences: A Dataset for Lifelong Place RecognitionFrederik Warburg, Søren Hauberg, Manuel López-Antequera, Pau Gargallo et al.CVPR 2020
- Thinking in Dynamics: How Multimodal Large Language Models Perceive, Track, and Reason Dynamics in Physical 4D WorldYuzhi Huang, Kairun Wen, Rongxin Gao, Dongxuan Liu et al.CVPR 2026 · 15 citations
- City-VLM: Towards Multidomain Perception Scene Understanding via Multimodal Incomplete LearningPenglei Sun, Yaoxian Song, Xiangru Zhu, Xiang Liu et al.ACM MM 2025 · 2 citations
- VSCD: Video-based Scene Change Detection in Unaligned ScenesJiae Yoon, Ue-Hwan KimICML 2026
- Seeing the Scene Matters: Revealing Forgetting in Video Understanding Models with a Scene-Aware Long-Video BenchmarkSeng Nam Chen, Hao Chen, Chenglam Ho, Xinyu Mao et al.CVPR 2026
