Efficient Representation Learning of Satellite Image Time Series and Their Fusion for Spatiotemporal Applications
Poonam Goyal, Arshveer Kaur, Arvind Ram, Navneet Goyal
摘要
Satellite data bolstered by their increasing accessibility is leading to many endeavors of automated monitoring of the earth's surface for various applications. Such applications demand high spatial resolution images at a temporal resolution of a few days which entails the challenge of processing a huge volume of image time series data. To overcome this computing bottleneck, we present PatchNet, a bespoke adaptation of beam search and attention mechanism. PatchNet is an automated patch selection neural network that requires only a partial spatial traversal of an image time series and yet achieves impressive results. Satellite systems face a trade-off between spatial and temporal resolutions due to budget/technical constraints e.g., Landsat-8/9 or Sentinel-2 have high spatial resolution whereas, MODIS has high temporal resolution. To deal with the limitation of coarse temporal resolution, we propose FuSITSNet, a twofold feature-based generic fusion model with multimodal learning in a contrastive setting. It produces a learned representation after fusion of two satellite image time series leveraging finer spatial resolution of Landsat and finer temporal resolution of MODIS. The patch alignment module of FuSITSNet aligns the PatchNet processed patches of Landsat-8 with the corresponding MODIS regions to incorporate its finer resolution temporal features. The untraversed patches are handled by the cross-modality attention which highlights additional hot spot features from the two modalities. We conduct extensive experiments on more than 2000 counties of US for crop yield, snow cover, and solar energy prediction and show that even one-fourth spatial processing of image time series produces state-of-the-art results. FuSITSNet outperforms the predictions of single modality and data obtained using existing generative fusion models and allows for monitoring of dynamic phenomena using freely accessible images, thereby unlocking new opportunities.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper2
相关 Paper
- Panoptic Segmentation of Satellite Image Time Series with Convolutional Temporal Attention NetworksVivien Sainte Fare Garnot, Loïc LandrieuICCV 2021 · 被引用 245 次
- PAN-Crafter: Learning Modality-Consistent Alignment for Pan-SharpeningJeonghyeok Do, Sungpyo Kim, Geunhyuk Youk, Jaehyup Lee 等ICCV 2025 · 被引用 3 次
- Semantic-Adaptive Diffusion for Dynamic Spatiotemporal FusionJinsong Zhang, Ying Qu, Yuan Liao, Hairong Qi 等CVPR 2026
- U2Net: A General Framework with Spatial-Spectral-Integrated Double U-Net for Image FusionSiran Peng, Chenhao Guo, Xiao Wu, Liang-Jian DengACM MM 2023 · 被引用 43 次
- MMST-ViT: Climate Change-aware Crop Yield Prediction via Multi-Modal Spatial-Temporal Vision TransformerFudong Lin, Summer Crawford, Kaleb Guillot, Yihe Zhang 等ICCV 2023 · 被引用 69 次
