Unifying Heterogeneous Multi-Modal Remote Sensing Detection Via Language-Pivoted Pretraining
Yuxuan Li, Yuming Chen, Yunheng Li, Ming-Ming Cheng, Xiang Li, jian Yang
摘要
Heterogeneous multi-modal remote sensing object detection aims to accurately detect objects from diverse sensors (e.g., RGB, SAR, Infrared). Existing approaches largely adopt a late alignment paradigm, in which modality alignment and task-specific optimization are entangled during downstream fine-tuning. This tight coupling complicates optimization and often results in unstable training and suboptimal generalization. To address these limitations, we propose BabelRS, a unified language-pivoted pretraining framework that explicitly decouples modality alignment from downstream task learning. BabelRS comprises two key components: Concept-Shared Instruction Aligning (CSIA) and Layerwise Visual-Semantic Annealing (LVSA). CSIA aligns each sensor modality to a shared set of linguistic concepts, using language as a semantic pivot to bridge heterogeneous visual representations. To further mitigate the granularity mismatch between high-level language representations and dense detection objectives, LVSA progressively aggregates multi-scale visual features to provide fine-grained semantic guidance. Extensive experiments demonstrate that BabelRS stabilizes training and consistently outperforms state-of-the-art methods without bells and whistles. Code: github.com/zcablii/SM3Det.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper19
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 被引用 11,349 次
- R3Det: Refined Single-Stage Detector with Feature Refinement for Rotating ObjectXue Yang, Junchi Yan, Ziming Feng, Tao HeAAAI 2021 · 被引用 1,109 次
- Large Selective Kernel Network for Remote Sensing Object DetectionYuxuan Li, Qibin Hou, Zhaohui Zheng, Ming-Ming Cheng 等ICCV 2023 · 被引用 535 次
- LanguageBind: Extending Video-Language Pretraining to N-modality by Language-based Semantic AlignmentBin Zhu, Bin Lin, Munan Ning, Yang Yan 等ICLR 2024 · 被引用 403 次
相关 Paper
- SM3Det: A Unified Model for Multi-Modal Remote Sensing Object DetectionYuxuan Li, Xiang Li, Yunheng Li, Yicheng Zhang 等AAAI 2026 · 被引用 25 次
- VLM4RSDet: Collaborative Optimization with Vision-Language Model for Enhancing Remote Sensing Object DetectionShuohao Shi, Qiang Fang, Xin XuCVPR 2026
- Vision-Language Pre-Training for Multimodal Aspect-Based Sentiment AnalysisYan Ling, Jianfei Yu, Rui XiaACL 2022 · 被引用 116 次
- AlignDet: Aligning Pre-training and Fine-tuning in Object DetectionMing Li, Jie Wu, Xionghui Wang, Chen Chen 等ICCV 2023 · 被引用 36 次
- MVPTR: Multi-Level Semantic Alignment for Vision-Language Pre-Training via Multi-Stage LearningZejun Li, Zhihao Fan, Huaixiao Tou, Jingjing Chen 等ACM MM 2022 · 被引用 15 次
