Towards Geospatial Foundation Models via Continual Pretraining
Matías Mendieta, Boran Han, Xingjian Shi, Yi Zhu, Chen Chen
摘要
Geospatial technologies are becoming increasingly essential in our world for a wide range of applications, including agriculture, urban planning, and disaster response. To help improve the applicability and performance of deep learning models on these geospatial tasks, various works have begun investigating foundation models for this domain. Researchers have explored two prominent approaches for introducing such models in geospatial applications, but both have drawbacks in terms of limited performance benefit or prohibitive training cost. Therefore, in this work, we propose a novel paradigm for building highly effective geospatial foundation models with minimal resource cost and carbon impact. We first construct a compact yet diverse dataset from multiple sources to promote feature diversity, which we term GeoPile. Then, we investigate the potential of continual pretraining from large-scale ImageNet-22k models and propose a multi-objective continual pretraining paradigm, which leverages the strong representations of ImageNet while simultaneously providing the freedom to learn valuable in-domain features. Our approach outperforms previous state-of-the-art geospatial pretraining methods in an extensive evaluation on seven downstream datasets covering various tasks such as change detection, classification, multi-label classification, semantic segmentation, and super-resolution. Code is available at https://github.com/mmendiet/GFM.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper33
- TESSERA: Temporal Embeddings of Surface Spectra for Earth Representation and AnalysisZhengpeng Feng, Clement Atzberger, Sadiq Jaffer, Jovana Knezevic 等CVPR 2026 · 被引用 61 次
- Rethinking Transformers Pre-training for Multi-Spectral Satellite ImageryMubashir Noman, Muzammal Naseer, Hisham Cholakkal, Rao Muhammad Anwer 等CVPR 2024 · 被引用 51 次
- D-CPT Law: Domain-specific Continual Pre-Training Scaling Law for Large Language ModelsHaoran Que, Jiaheng Liu, Ge Zhang, Chenchen Zhang 等NeurIPS 2024 · 被引用 47 次
- TerraMind: Large-Scale Generative Multimodality for Earth ObservationJohannes Jakubik, Felix Yang, Benedikt Blumenstiel, Erik Scheurer 等ICCV 2025 · 被引用 43 次
- ReFound: Crafting a Foundation Model for Urban Region Understanding upon Language and Visual FoundationsCongxi Xiao, Jingbo Zhou, Yixiong Xiao, Jizhou Huang 等KDD 2024 · 被引用 18 次
它引用的顶会 Paper12
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu 等ICCV 2021 · 被引用 31,683 次
- Bootstrap Your Own Latent - A New Approach to Self-Supervised LearningJean-Bastien Grill, Florian Strub, Florent Altché, Corentin Tallec 等NeurIPS 2020 · 被引用 9,171 次
- BEiT: BERT Pre-Training of Image TransformersHangbo Bao, Li Dong, Songhao Piao, Furu WeiICLR 2022 · 被引用 3,632 次
- Generative Pretraining From PixelsMark Chen, Alec Radford, Rewon Child, Jeffrey Wu 等ICML 2020 · 被引用 1,773 次
- SimMIM: a Simple Framework for Masked Image ModelingZhenda Xie, Zheng Zhang, Yue Cao, Yutong Lin 等CVPR 2022 · 被引用 1,129 次
相关 Paper
- GeoSANE: Learning Geospatial Representations from Models, Not DataJoëlle Hanna, Damian Falk, Stella X. Yu, Damian BorthCVPR 2026 · 被引用 1 次
- PhySwin: An Efficient and Physically-Informed Foundation Model for Multispectral Earth ObservationChong Tang, Joseph Powell, Dirk Koch, Robert Mullins 等NeurIPS 2025 · 被引用 1 次
- Parameter-Efficient Adaptation of Geospatial Foundation Models Through Embedding DeflectionRomain Thoreau, Valerio Marsocci, Dawa DerksenICCV 2025 · 被引用 1 次
- Can Generative Geospatial Diffusion Models Excel as Discriminative Geospatial Foundation Models?Yuru Jia, Valerio Marsocci, Ziyang Gong, Xue Yang 等ICCV 2025
- Bridging Remote Sensors with Multisensor Geospatial Foundation ModelsBoran Han, Shuai Zhang, Xingjian Shi, Markus ReichsteinCVPR 2024
