SM3Det: A Unified Model for Multi-Modal Remote Sensing Object Detection
Yuxuan Li, Xiang Li, Yunheng Li, Yicheng Zhang, Yimian Dai, Qibin Hou, Ming-Ming Cheng, Jian Yang
Abstract
With the rapid advancement of remote sensing technology, high-resolution multi-modal imagery is now more widely accessible. Conventional object detection models are trained on a single dataset, often restricted to a specific imaging modality and annotation format. However, such an approach overlooks the valuable shared knowledge across multi-modalities and limits the model’s applicability in more versatile scenarios. This paper introduces a new task called Multi-Modal Datasets and Multi-Task Object Detection (M2Det) for remote sensing, designed to accurately detect horizontal or oriented objects from any sensor modality. This task poses challenges due to 1) the trade-offs involved in managing multi-modal modelling and 2) the complexities of multi-task optimization. To address these, we establish a benchmark dataset and propose a unified model, SM3Det (Single Model for Multi-Modal datasets and Multi-Task object Detection). SM3Det leverages a grid-level sparse MoE backbone to enable joint knowledge learning while preserving distinct feature representations for different modalities. Furthermore, we propose a novel consistency and synchronization optimization mechanism, allowing it to effectively handle varying levels of learning difficulty across modalities and tasks. Extensive experiments demonstrate SM3Det's effectiveness and generalizability, consistently outperforming the combination of specialized models on individual datasets.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 4db9c363-6ba8-42aa-811b-7980d4404a55Cited by top-tier papers11
- Earth-Agent: Unlocking the Full Landscape of Earth Observation with AgentsPeilin Feng, Zhutao Lv, Junyan Ye, Xiaolei Wang et al.ICLR 2026 · 49 citations
- Strip R-CNN: Large Strip Convolution for Remote Sensing Object DetectionXinbin Yuan, Zhaohui Zheng, Yuxuan Li, Xialei Liu et al.AAAI 2026 · 32 citations
- NAIPv2: Debiased Pairwise Learning for Efficient Paper Quality EstimationPenghai Zhao, Jinyu Tian, Qinghua Xing, Xin Zhang et al.ICLR 2026 · 6 citations
- DenoDet V2: Phase-Amplitude Cross Denoising for SAR Object DetectionKang Ni, Minrui Zou, Yuxuan Li, Xiang Li et al.AAAI 2026 · 2 citations
- DISTA-Net: Dynamic Closely-Spaced Infrared Small Target UnmixingShengdong Han, Shangdong Yang, Yuxuan Li, Xin Zhang et al.ICCV 2025 · 2 citations
Builds on14
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- A ConvNet for the 2020sZhuang Liu, Hanzi Mao, Chao-Yuan Wu, Christoph Feichtenhofer et al.CVPR 2022 · 6,782 citations
- Scaling Vision with Sparse Mixture of ExpertsCarlos Riquelme, Joan Puigcerver, Basil Mustafa, Maxim Neumann et al.NeurIPS 2021 · 1,213 citations
- Rethinking Rotated Object Detection with Gaussian Wasserstein Distance LossXue Yang, Junchi Yan, Qi Ming, Wentao Wang et al.ICML 2021 · 572 citations
- Large Selective Kernel Network for Remote Sensing Object DetectionYuxuan Li, Qibin Hou, Zhaohui Zheng, Ming-Ming Cheng et al.ICCV 2023 · 535 citations
Related papers
- Unifying Heterogeneous Multi-Modal Remote Sensing Detection Via Language-Pivoted PretrainingYuxuan Li, Yuming Chen, Yunheng Li, Ming-Ming Cheng et al.ICML 2026
- MODA: The First Challenging Benchmark for Multispectral Object Detection in Aerial ImagesShuaihao Han, Tingfa Xu, Peifu Liu, Jianan LiAAAI 2026 · 1 citation
- OpenRSD: Towards Open-Prompts for Object Detection in Remote Sensing ImagesZiyue Huang, Yongchao Feng, Ziqi Liu, Shuai Yang et al.ICCV 2025 · 3 citations
- Sparse-to-dense Multimodal Image Registration via Multi-Task LearningKaining Zhang, Jiayi MaICML 2024 · 6 citations
- Serial Over Parallel: Learning Continual Unification for Multi-Modal Visual Object Tracking and BenchmarkingZhangyong Tang, Tianyang Xu, Xuefeng Zhu, Chunyang Cheng et al.ACM MM 2025 · 2 citations
