Change-UP: Advancing Visualization and Inference Capability for Multi-level Remote Sensing Change Interpretation
Mo Yang, Luo Chen, Jiali Zhou
Abstract
Monitoring Earth's surface evolution is critical for understanding environmental processes and anthropogenic impacts, necessitating precise analysis of multi-temporal remote sensing data via integrated change detection and semantic interpretation. While existing methods achieve either pixel-level localization or semantic-level description, their conventional isolated frameworks exhibit two critical limitations: 1) inability to establish multimodal semantic consistency, and 2) feature confusion under complex environmental noise. To address these challenges, we propose Change-UP, a unified learning framework that synergizes Multilevel Change Interpretation (MCI) with Large Language Models (LLMs), enabling comprehensive analysis through dual visualization-inference perspectives. The architecture comprises three key components: 1) Adaptive Adjustment Change-Aware (A2CA) captures multi-scale inter-temporal differences through pyramidal feature fusion, significantly enhancing signal-to-noise ratio of differential features. 2) Query Semantic Consistency (QSC) establishes cross-modal alignment through learnable semantic queries, improving feature discrimination on challenging samples. 3) Style Transform Diversity (STD) integrates text-driven style transform with style synergy optimization to enhance inter-modal diversity while preventing feature space collapse. Extensive experiments demonstrate state-of-the-art performance, achieving 87.02% MIoU and CIDEr-D 143.40, surpassing previous methods by 0.59% and 3.11%, respectively. The framework's novel integration of visual-linguistic modalities opens new possibilities for intelligent Earth observation systems.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Related papers
- UniChange: Unifying Change Detection with Multimodal Large Language ModelXu Zhang, Danyang Li, Xiaohang Dong, Tianhao Wu et al.CVPR 2026 · 9 citations
- Leveraging Textual Compositional Reasoning for Robust Change CaptioningKyu Ri Park, Jiyoung Park, Seong Tae Kim, Hong Joo Lee et al.AAAI 2026
- Change3D: Revisiting Change Detection and Captioning from A Video Modeling PerspectiveDuowang Zhu, Xiaohu Huang, Haiyan Huang, Hao Zhou et al.CVPR 2025
- Hugging Visual Prompt and Segmentation Tokens: Consistency Learning for Fine-Grained Visual Understanding in MLLMsjing yang, Sen Yang, Boqiang Duan, Ming Dai et al.CVPR 2026
- TerraScope: Pixel-Grounded Visual Reasoning for Earth ObservationYan Shu, Bin Ren, Zhitong Xiong, Xiao Xiang Zhu et al.CVPR 2026 · 9 citations
