ROS-SAM: High-Quality Interactive Segmentation for Remote Sensing Moving Object
Zhe Shan, Yang Liu, Lei Zhou, Cheng Yan, Heng Wang, Xia Xie
Abstract
The availability of large-scale remote sensing video data underscores the importance of high-quality interactive segmentation. However, challenges such as small object sizes, ambiguous features, and limited generalization make it difficult for current methods to achieve this goal. In this work, we propose ROS-SAM, a method designed to achieve high-quality interactive segmentation while preserving generalization across diverse remote sensing data. The ROS-SAM is built upon three key innovations: 1) LoRAbased fine-tuning, which enables efficient domain adaptation while maintaining SAM's generalization ability, 2) Enhancement of network deep layers to improve the discriminability of extracted features, thereby reducing misclassifications, and 3) Integration of global context with local boundary details in the mask decoder to generate highquality segmentation masks. Additionally, we redesign the data pipeline to ensure the model learns to better handle objects at varying scales during training while focusing on high-quality predictions during inference. Experiments on remote sensing video datasets show that the data pipeline boosts the IoU by 6%, while ROS-SAM increases the IoU by 13%. Finally, when evaluated on existing remote sensing object tracking datasets, ROS-SAM demonstrates impressive zero-shot capabilities, generating masks that closely resemble manual annotations. These results confirm ROS-SAM as a powerful tool for fine-grained segmentation in remote sensing applications. Code is available at: https://github.com/ShanZard/ROS-SAM .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext f2522e4f-1ab2-48a6-9619-0f6e51b5b8dfCited by top-tier papers5
- InstructSAM: A Training-free Framework for Instruction-Oriented Remote Sensing Object RecognitionYijie Zheng, Weijie Wu, Qingyun Li, Xuehui Wang et al.NeurIPS 2025 · 12 citations
- RSVG-ZeroOV: Exploring a Training-Free Framework for Zero-Shot Open-Vocabulary Visual Grounding in Remote Sensing ImagesKe Li, Di Wang, Ting Wang, Fuyu Dong et al.AAAI 2026 · 7 citations
- ReSAM: Refine, Requery, and Reinforce: Self-Prompting Point-Supervised Segmentation for Remote Sensing ImagesMuhammad Naseer SubhaniCVPR 2026 · 2 citations
- FICGen: Frequency-Inspired Contextual Disentanglement for Layout-driven Degraded Image GenerationWenzhuang Wang, Yifan Zhao, Mingcan Ma, Ming Liu et al.ICCV 2025 · 1 citation
- CrossCut: Cross-Patch Aware Interactive Segmentation for Remote Sensing ImagesZheng Lin, Nan Zhou, Yuhan Wang, Bojian ZhangAAAI 2026
Builds on15
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao et al.ICCV 2023 · 13,211 citations
- Segment Anything in High QualityLei Ke, Mingqiao Ye, Martin Danelljan, Yifan Liu et al.NeurIPS 2023 · 709 citations
- Seasonal Contrast: Unsupervised Pre-Training from Uncurated Remote Sensing DataOscar Mañas, Alexandre Lacoste, Xavier Giró-i-Nieto, David Vázquez et al.ICCV 2021 · 361 citations
Related papers
- Convolution Meets LoRA: Parameter Efficient Finetuning for Segment Anything ModelZihan Zhong, Zhiqiang Tang, Tong He, Haoyang Fang et al.ICLR 2024 · 91 citations
- Towards Fine-Grained Interactive Segmentation in Images and VideosYuan Yao, Qiushi Yang, Miaomiao Cui, Liefeng BoICCV 2025 · 2 citations
- MTSAM: Multi-Task Fine-Tuning for Segment Anything ModelXuehao Wang, Zhan Zhuang, Feiyang Ye, Yu ZhangICLR 2025
- Live Interactive Training for Video SegmentationXinyu Yang, Haozheng Yu, Yihong Sun, Bharath Hariharan et al.CVPR 2026 · 1 citation
- Correspondence as Video: Test-Time Adaption on SAM2 for Reference Segmentation in the WildHaoran Wang, Zekun Li, Jian Zhang, Lei Qi et al.ICCV 2025
