VISO: Accelerating In-Orbit Object Detection with Language-Guided Mask Learning and Sparse Inference
Meiqi Wang, Han Qiu
Abstract
In-orbit object detection is essential for Earth observation missions on satellites equipped with GPUs. A promising approach is to use pre-trained vision-language modeling (VLM) to enhance its open-vocabulary capability. However, adopting it on satellites poses two challenges: (1) satellite imagery differs substantially from natural images, and (2) satellites' embedded GPUs are insufficient for complex models' inference. We reveal their lack of a crucial prior: in-orbit detection involves identifying a set of known objects within a cluttered yet monotonous background. Motivated by this observation, we propose VISO, a Visionlanguage Instructed Satellite Object detection model that focuses on object-specific features while suppressing irrelevant regions through language-guided mask learning. After pre-training on a large-scale satellite dataset with 3.4M region-text pairs, VISO enhances object-text alignment and object-centric features to improve detection accuracy. Also, VISO suppresses irrelevant regions, enabling highly sparse inference to accelerate speed on satellites. Extensive experiments show that VISO without sparsity outperforms stateof-the-art (SOTA) VLMs in zero-shot detection by increasing 34.1% AP and reducing 27× FLOPs, and surpasses specialist models in supervised object detection and object referring by improving 2.3% AP. When sparsifying VISO to a comparable AP, FLOPs can be greatly reduced by up to 8.5×. Real-world tests reveal that VISO achieves a 2.8-4.8× FPS speed-up on satellites' embedded GPUs 1 .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on14
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- Generalized Focal Loss: Learning Qualified and Distributed Bounding Boxes for Dense Object DetectionXiang Li, Wenhai Wang, Lijun Wu, Shuo Chen et al.NeurIPS 2020 · 2,118 citations
- Large Selective Kernel Network for Remote Sensing Object DetectionYuxuan Li, Qibin Hou, Zhaohui Zheng, Ming-Ming Cheng et al.ICCV 2023 · 535 citations
- QueryDet: Cascaded Sparse Query for Accelerating High-Resolution Small Object DetectionChenhongyi Yang, Zehao Huang, Naiyan WangCVPR 2022 · 472 citations
Related papers
- YOLO-World: Real-Time Open-Vocabulary Object DetectionTianheng Cheng, Lin Song, Yixiao Ge, Wenyu Liu et al.CVPR 2024
- Aligning Bag of Regions for Open-Vocabulary Object DetectionSize Wu, Wenwei Zhang, Sheng Jin, Wentao Liu et al.CVPR 2023
- Learning to Prompt for Open-Vocabulary Object Detection with Vision-Language ModelYu Du, Fangyun Wei, Zihe Zhang, Miaojing Shi et al.CVPR 2022 · 311 citations
- Remote Sensing Vision-Language Foundation Models without Annotations via Ground Remote AlignmentUtkarsh Mall, Cheng Perng Phoo, Meilin Kelsey Liu, Carl Vondrick et al.ICLR 2024 · 90 citations
- VLM4RSDet: Collaborative Optimization with Vision-Language Model for Enhancing Remote Sensing Object DetectionShuohao Shi, Qiang Fang, Xin XuCVPR 2026
