Semantic-Adaptive Diffusion for Dynamic Spatiotemporal Fusion
Jinsong Zhang, Ying Qu, Yuan Liao, Hairong Qi, Zhenzhou Shao
Abstract
Frequent and precise land surface monitoring is critical for numerous applications, yet existing satellites struggle to achieve both simultaneously. Spatiotemporal fusion (STF) tackles this challenge by integrating multiple satellite images to generate data with improved temporal and spatial resolution, enabling more frequent and precise land surface observations. However, current methods often fail to recover dynamic landscape changes due to significant scale discrepancies among multi-source images. To address these challenges, we propose a semantic-adaptive diffusion framework for dynamic spatiotemporal fusion (SA-STF), which constrains the solution space using low-resolution and high-frequency features decoupled via a Taylor-inspired decoder. By incorporating temporal feature alignment and semantic-adaptive fusion modules, SA-STF projects multimodal images with temporal dynamics into a unified latent space, and adaptively enhances spatial details while maintaining the spectral consistency of the reconstructed images. Experiments on benchmark datasets demonstrate that SA-STF outperforms existing methods in both quantitative and qualitative evaluations, particularly in complex and dynamic scenes.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on3
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- Directly Denoising Diffusion ModelsDan Zhang, Jingjing Wang, Feng LuoICML 2024 · 11,724 citations
- Residual Denoising Diffusion ModelsJiawei Liu, Qiang Wang, Huijie Fan, Yinong Wang et al.CVPR 2024 · 96 citations
Related papers
- Dynamic View Synthesis with Spatio-Temporal Feature Warping from Sparse ViewsDeqi Li, Shi-Sheng Huang, Tianyu Shen, Hua HuangACM MM 2023 · 6 citations
- Efficient Representation Learning of Satellite Image Time Series and Their Fusion for Spatiotemporal ApplicationsPoonam Goyal, Arshveer Kaur, Arvind Ram, Navneet GoyalAAAI 2024 · 3 citations
- Learning Spatial Adaptation and Temporal Coherence in Diffusion Models for Video Super-ResolutionZhikai Chen, Fuchen Long, Zhaofan Qiu, Ting Yao et al.CVPR 2024
- Satellite to GroundScape - Large-scale Consistent Ground View Generation from Satellite ViewsNingli Xu, Rongjun QinCVPR 2025
- PAN-Crafter: Learning Modality-Consistent Alignment for Pan-SharpeningJeonghyeok Do, Sungpyo Kim, Geunhyuk Youk, Jaehyup Lee et al.ICCV 2025 · 3 citations
