Brewing Stronger Features: Dual-Teacher Distillation for Multispectral Earth Observation
Filip Wolf, Blaz Rolih, Luka Cehovin Zajc
摘要
Foundation models are transforming Earth Observation (EO), yet the diversity of EO sensors and modalities makes a single universal model unrealistic. Multiple specialized EO foundation models (EOFMs) will likely coexist, making efficient knowledge transfer across modalities essential. Most existing EO pretraining relies on masked image modeling, which emphasizes local reconstruction but provides limited control over global semantic structure. To address this, we propose a dual-teacher contrastive distillation framework for multispectral imagery that aligns the student's pretraining objective with the contrastive self-distillation paradigm of modern optical vision foundation models (VFMs). Our approach combines a multispectral teacher with an optical VFM teacher, enabling coherent cross-modal representation learning. Experiments across diverse optical and multispectral benchmarks show that our model adapts to multispectral data without compromising performance on optical-only inputs, achieving state-of-the-art results in both settings, with an average improvement of 3.64 percentage points in semantic segmentation, 1.2 in change detection, and 1.31 in classification tasks. This demonstrates that contrastive distillation provides a principled and efficient approach to scalable representation learning across heterogeneous EO data sources. Project page: magentahttps://wolfilip.github.io/DEO/.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper31
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu 等ICCV 2021 · 被引用 31,683 次
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 被引用 24,064 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao 等ICCV 2023 · 被引用 13,211 次
相关 Paper
- Fine-grained Image-to-LiDAR Contrastive Distillation with Visual Foundation ModelsYifan Zhang, Junhui HouNeurIPS 2024 · 被引用 9 次
- Generalizable Knowledge Distillation from Vision Foundation Models for Semantic SegmentationChonghua Lv, Dong Zhao, Shuang Wang, Dou Quan 等CVPR 2026 · 被引用 1 次
- A Mixed Diet Makes DINO An Omnivorous Vision EncoderRishabh Kabra, Maks Ovsjanikov, Drew A. Hudson, Ye Xia 等CVPR 2026 · 被引用 3 次
- SGPFeat: Semantic and Geometric Priors for Multi-modal Image MatchingYuxin Deng, Botian Wang, Kaining Zhang, Hao Zhang 等AAAI 2026
- PRISM: Synergizing Vision Foundation Models via Self-organized Expert SpecializationYing Tang, Dong Li, Youjia Zhang, Zikai Song 等ICML 2026
