SM3Det: A Unified Model for Multi-Modal Remote Sensing Object Detection
Yuxuan Li, Xiang Li, Yunheng Li, Yicheng Zhang, Yimian Dai, Qibin Hou, Ming-Ming Cheng, Jian Yang
摘要
With the rapid advancement of remote sensing technology, high-resolution multi-modal imagery is now more widely accessible. Conventional object detection models are trained on a single dataset, often restricted to a specific imaging modality and annotation format. However, such an approach overlooks the valuable shared knowledge across multi-modalities and limits the model’s applicability in more versatile scenarios. This paper introduces a new task called Multi-Modal Datasets and Multi-Task Object Detection (M2Det) for remote sensing, designed to accurately detect horizontal or oriented objects from any sensor modality. This task poses challenges due to 1) the trade-offs involved in managing multi-modal modelling and 2) the complexities of multi-task optimization. To address these, we establish a benchmark dataset and propose a unified model, SM3Det (Single Model for Multi-Modal datasets and Multi-Task object Detection). SM3Det leverages a grid-level sparse MoE backbone to enable joint knowledge learning while preserving distinct feature representations for different modalities. Furthermore, we propose a novel consistency and synchronization optimization mechanism, allowing it to effectively handle varying levels of learning difficulty across modalities and tasks. Extensive experiments demonstrate SM3Det's effectiveness and generalizability, consistently outperforming the combination of specialized models on individual datasets.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper11
- Earth-Agent: Unlocking the Full Landscape of Earth Observation with AgentsPeilin Feng, Zhutao Lv, Junyan Ye, Xiaolei Wang 等ICLR 2026 · 被引用 49 次
- Strip R-CNN: Large Strip Convolution for Remote Sensing Object DetectionXinbin Yuan, Zhaohui Zheng, Yuxuan Li, Xialei Liu 等AAAI 2026 · 被引用 32 次
- NAIPv2: Debiased Pairwise Learning for Efficient Paper Quality EstimationPenghai Zhao, Jinyu Tian, Qinghua Xing, Xin Zhang 等ICLR 2026 · 被引用 6 次
- DenoDet V2: Phase-Amplitude Cross Denoising for SAR Object DetectionKang Ni, Minrui Zou, Yuxuan Li, Xiang Li 等AAAI 2026 · 被引用 2 次
- DISTA-Net: Dynamic Closely-Spaced Infrared Small Target UnmixingShengdong Han, Shangdong Yang, Yuxuan Li, Xin Zhang 等ICCV 2025 · 被引用 2 次
它引用的顶会 Paper14
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu 等ICCV 2021 · 被引用 31,683 次
- A ConvNet for the 2020sZhuang Liu, Hanzi Mao, Chao-Yuan Wu, Christoph Feichtenhofer 等CVPR 2022 · 被引用 6,782 次
- Scaling Vision with Sparse Mixture of ExpertsCarlos Riquelme, Joan Puigcerver, Basil Mustafa, Maxim Neumann 等NeurIPS 2021 · 被引用 1,213 次
- Rethinking Rotated Object Detection with Gaussian Wasserstein Distance LossXue Yang, Junchi Yan, Qi Ming, Wentao Wang 等ICML 2021 · 被引用 572 次
- Large Selective Kernel Network for Remote Sensing Object DetectionYuxuan Li, Qibin Hou, Zhaohui Zheng, Ming-Ming Cheng 等ICCV 2023 · 被引用 535 次
相关 Paper
- Unifying Heterogeneous Multi-Modal Remote Sensing Detection Via Language-Pivoted PretrainingYuxuan Li, Yuming Chen, Yunheng Li, Ming-Ming Cheng 等ICML 2026
- MODA: The First Challenging Benchmark for Multispectral Object Detection in Aerial ImagesShuaihao Han, Tingfa Xu, Peifu Liu, Jianan LiAAAI 2026 · 被引用 1 次
- OpenRSD: Towards Open-Prompts for Object Detection in Remote Sensing ImagesZiyue Huang, Yongchao Feng, Ziqi Liu, Shuai Yang 等ICCV 2025 · 被引用 3 次
- Sparse-to-dense Multimodal Image Registration via Multi-Task LearningKaining Zhang, Jiayi MaICML 2024 · 被引用 6 次
- Serial Over Parallel: Learning Continual Unification for Multi-Modal Visual Object Tracking and BenchmarkingZhangyong Tang, Tianyang Xu, Xuefeng Zhu, Chunyang Cheng 等ACM MM 2025 · 被引用 2 次
