LDS2AE: Local Diffusion Shared-Specific Autoencoder for Multimodal Remote Sensing Image Classification with Arbitrary Missing Modalities
Jiahui Qu, Yuanbo Yang, Wenqian Dong, Yufei Yang
Abstract
Recent research on the joint classification of multimodal remote sensing data has achieved great success. However, due to the limitations imposed by imaging conditions, the case of missing modalities often occurs in practice. Most previous researchers regard the classification in case of different missing modalities as independent tasks. They train a specific classification model for each fixed missing modality by extracting multimodal joint representation, which cannot handle the classification of arbitrary (including multiple and random) missing modalities. In this work, we propose a local diffusion shared-specific autoencoder (LDS 2 AE), which solves the classification of arbitrary missing modalities with a single model. The LDS 2 AE captures the data distribution of different modalities to learn multimodal shared feature for classification by designing a novel local diffusion autoencoder which consists of a modality-shared encoder and several modality-specific decoders. The modality-shared encoder is designed to extract multimodal shared feature by employing the same parameters to map multimodal data into a shared subspace. The modality-specific decoders put the multimodal shared feature to reconstruct the image of each modality, which facilitates the shared feature to learn unique information of different modalities. In addition, we incorporate masked training to the diffusion autoencoder to achieve local diffusion, which significantly reduces the training cost of model. The approach is tested on widely-used multimodal remote sensing datasets, demonstrating the effectiveness of the proposed LDS 2 AE in addressing the classification of arbitrary missing modalities. The code is available at https://github.com/Jiahuiqu/LDS2AE.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers4
- DiffCLIP: Few-shot Language-driven Multimodal ClassifierJiaqing Zhang, Mingxiang Cao, Xue Yang, Kai Jiang et al.AAAI 2025 · 4 citations
- Deep Fuzzy Multi-view Learning for Reliable ClassificationSiyuan Duan, Yuan Sun, Dezhong Peng, Guiduo Duan et al.ICML 2025
- Can Generative Geospatial Diffusion Models Excel as Discriminative Geospatial Foundation Models?Yuru Jia, Valerio Marsocci, Ziyang Gong, Xue Yang et al.ICCV 2025
- Frequency-Aware Vision-Language Multimodality Generalization Network for Remote Sensing Image ClassificationJunjie Zhang, Feng Zhao, Hanqiang Liu, Jun YuAAAI 2026
Builds on7
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 13,211 citations
- Improved Denoising Diffusion Probabilistic ModelsAlexander Quinn Nichol, Prafulla DhariwalICML 2021 · 5,234 citations
- BEiT: BERT Pre-Training of Image TransformersHangbo Bao, Li Dong, Songhao Piao, Furu WeiICLR 2022 · 3,632 citations
- MISA: Modality-Invariant and -Specific Representations for Multimodal Sentiment AnalysisDevamanyu Hazarika, Roger Zimmermann, Soujanya PoriaACM MM 2020 · 1,037 citations
Related papers
- Language-Guided Visual Prompt Compensation for Multi-Modal Remote Sensing Image Classification with Modality AbsenceLing Huang, Wenqian Dong, Song Xiao, Jiahui Qu et al.ACM MM 2024 · 1 citation
- M3AE: Multimodal Representation Learning for Brain Tumor Segmentation with Missing ModalitiesHong Liu, Dong Wei, Donghuan Lu, Jinghan Sun et al.AAAI 2023 · 101 citations
- Auto-GAN: Self-Supervised Collaborative Learning for Medical Image SynthesisBing Cao, Han Zhang, Nannan Wang, Xinbo Gao et al.AAAI 2020 · 94 citations
- Multi-Modal Learning with Missing Modality via Shared-Specific Feature ModellingHu Wang, Yuanhong Chen, Congbo Ma, Jodie Avery et al.CVPR 2023
- Masked-Diffusion Autoencoders for 3D Medical Vision Representation LearningJiachen Tu, Guanghui Qin, Theodore Zhengde Zhao, Jeya Maria Jose Valanarasu et al.CVPR 2026
