Spectrally Distilled Representations Aligned with Instruction-Augmented LLMs for Satellite Imagery
Minh Kha Do, Wei Xiang, Kang Han, Di Wu, Khoa T. Phan, Yi-Ping Phoebe Chen, Gaowen Liu, Ramana Kompella
Abstract
Vision-language foundation models (VLFMs) promise zero-shot and retrieval understanding for Earth observation. While operational satellite systems often lack full multi-spectral coverage, making RGB-only inference highly desirable for scalable deployment, the adoption of VLFMs for satellite imagery remains hindered by two factors: (1) multi-spectral inputs are informative but difficult to exploit consistently due to band redundancy and misalignment; and (2) CLIP-style text encoders limit semantic expressiveness and weaken fine-grained alignment. We present SATtxt, a spectrum-aware VLFM that operates with RGB inputs only at inference while retaining spectral cues learned during training. Our framework comprises two stages. First, Spectral Representation Distillation transfers spectral priors from a frozen multi-spectral teacher to an RGB student via a lightweight projector. Second, Spectrally Grounded Alignment with Instruction-Augmented LLMs bridges the distilled visual space and an expressive LLM embedding space. Across EuroSAT, BigEarthNet, and ForestNet, SATtxt improves zero-shot classification on average by 4.2%, retrieval by 5.9%, and linear probing by 2.7% over state-of-the-art baselines, demonstrating an efficient and deployable path toward spectrum-aware vision-language learning for Earth observation.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on19
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou et al.ICCV 2021 · 8,921 citations
- SatMAE: Pre-training Transformers for Temporal and Multi-Spectral Satellite ImageryYezhen Cong, Samar Khanna, Chenlin Meng, Patrick Liu et al.NeurIPS 2022 · 707 citations
- Scale-MAE: A Scale-Aware Masked Autoencoder for Multiscale Geospatial Representation LearningColorado J. Reed, Ritwik Gupta, Shufan Li, Sarah Brockman et al.ICCV 2023 · 373 citations
- Seasonal Contrast: Unsupervised Pre-Training from Uncurated Remote Sensing DataOscar Mañas, Alexandre Lacoste, Xavier Giró-i-Nieto, David Vázquez et al.ICCV 2021 · 361 citations
Related papers
- Building Vision-Language Models on Solid Foundations with Masked DistillationSepehr Sameni, Kushal Kafle, Hao Tan, Simon JenniCVPR 2024 · 4 citations
- AVION: Aerial Vision–Language Instruction from Offline Teacher to Prompt-Tuned NetworkYu Hu, Jianyang Gu, Hao Liu, Yue Cao et al.CVPR 2026
- S^2-KD: Semantic-Spectral Knowledge Distillation Spatiotemporal ForecastingWenshuo Wang, Yaomin Shen, Yingjie Tan, Yihao ChenAAAI 2026 · 5 citations
- Brewing Stronger Features: Dual-Teacher Distillation for Multispectral Earth ObservationFilip Wolf, Blaz Rolih, Luka Cehovin ZajcCVPR 2026 · 4 citations
- Distilling Vision-Language Foundation Models: A Data-Free Approach via Prompt DiversificationYunyi Xuan, Weijie Chen, Shicai Yang, Di Xie et al.ACM MM 2023 · 4 citations
