Deep Multi-Level Contrastive Clustering for Multi-Modal Remote Sensing Images
Weiqi Liu, Yongshan Zhang, Xinxin Wang, Lefei Zhang
Abstract
Multi-modal remote sensing image clustering aims to group similar pixels into the same cluster and separate dissimilar ones by leveraging the consistency and complementary information across multiple modalities, without relying on label guidance. Most existing deep learning-based methods address this task through a two-stage pipeline of feature learning followed by clustering, or adopt simple instance-level contrastive learning frameworks. In this paper, we propose an end-to-end deep multi-level contrastive clustering (DMLCC) model for multi-modal remote sensing images. The proposed DMLCC consists of three key components. Specifically, modality-specific vision encoders are initially employed to extract preliminary feature representations tailored to the characteristics of each modality. Spatial-spectral cross-modal fusion is then performed by integrating dedicated spatial and spectral feature extractors alongside a cross-modal fusion block to effectively capture and align complementary information across modalities. Finally, multi-level contrastive learning is applied to enhance feature discriminability through instance-level contrastive learning, while simultaneously promoting cluster separability via cluster-level contrastive learning. The network is trained in an end-to-end manner by integrating the three components to directly yield the clustering results. Extensive experiments on three datasets demonstrate that the proposed DMLCC outperforms state-of-the-art methods. Our code is publicly available at https://github.com/ZhangYongshan/DMLCC.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Cited by top-tier papers3
- SOAR: Semi-Supervised Open-Vocabulary Aerial Object Detection via Dual-Aware Enhanced Prior DenoisingXu Liu, Yihong Huang, Dan Zhang, Lingling Li et al.AAAI 2026
- Federated Multi-view Clustering for Remote Sensing DataRenxiang Guan, Xiang Yang, Hao Yu, Siwei Wang et al.ICML 2026
- Frequency-Aligned Cross-Modal Learning with Top-K Wavelet Fusion and Dynamic Expert Routing for Enhanced Retinal Disease DiagnosisYuxin Lin, Haoran Li, Haoyu Cao, Yongting Hu et al.AAAI 2026
Related papers
- EMVCC: Enhanced Multi-View Contrastive Clustering for Hyperspectral ImagesFulin Luo, Yi Liu, Xiuwen Gong, Zhixiong Nan et al.ACM MM 2024 · 21 citations
- Cross-View Progressive Feature Filtering for Multi-View Graph Clustering in Remote SensingBowen Liu, Xin Peng, Wenxuan Tu, Chengyao Wei et al.AAAI 2026
- Multi-view Graph Clustering with Dual Structure Awareness for Remote Sensing DataXin Peng, Bowen Liu, Renxiang Guan, Wenxuan TuACM MM 2025 · 4 citations
- End-to-End Adversarial-Attention Network for Multi-Modal ClusteringRunwu Zhou, Yi-Dong ShenCVPR 2020
- Cross-Domain Contrastive Learning for Time Series ClusteringFurong Peng, Jiachen Luo, Xuan Lu, Sheng Wang et al.AAAI 2024 · 13 citations
