RS2-SAM2: Customized SAM2 for Referring Remote Sensing Image Segmentation
Fu Rong, Meng Lan, Qian Zhang, Lefei Zhang
摘要
Referring Remote Sensing Image Segmentation (RRSIS) aims to segment target objects in remote sensing (RS) images based on textual descriptions. Although Segment Anything Model 2 (SAM2) has shown remarkable performance in various segmentation tasks, its application to RRSIS presents several challenges, including understanding the text-described RS scenes and generating effective prompts from text. To address these issues, we propose RS2-SAM2, a novel framework that adapts SAM2 to RRSIS by aligning the adapted RS features and textual features while providing pseudo-mask-based dense prompts. Specifically, we employ a union encoder to jointly encode the visual and textual inputs, generating aligned visual and text embeddings as well as multimodal class tokens. A bidirectional hierarchical fusion module is introduced to adapt SAM2 to RS scenes and align adapted visual features with the visually enhanced text embeddings, improving the model's interpretation of text-described RS scenes. To provide precise target cues for SAM2, we design a mask prompt generator, which takes the visual embeddings and class tokens as input and produces a pseudo-mask as the dense prompt of SAM2. Experimental results on several RRSIS benchmarks demonstrate that RS2-SAM2 achieves state-of-the-art performance.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- MPG-SAM 2: Adapting SAM 2 with Mask Priors and Global Context for Referring Video Object SegmentationFu Rong, Meng Lan, Qian Zhang, Lefei ZhangICCV 2025 · 被引用 4 次
- Any2RSI: Controllable Remote Sensing Text-to-Image Generation via Any Control and Enriched DescriptionXu Zhang, Jianzhong Huang, Lefei ZhangAAAI 2026 · 被引用 1 次
- ORSATR-X: A Foundation Model based on Differential-and-Excitation Networks for Optical Remote Sensing Object RecognitionCanyu Mo, Yongxiang Liu, Jiehua Zhang, Zilong Yu 等CVPR 2026 · 被引用 1 次
它引用的顶会 Paper19
- Unsupervised Cross-lingual Representation Learning at ScaleAlexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary 等ACL 2020 · 被引用 539 次
- Hiera: A Hierarchical Vision Transformer without the Bells-and-WhistlesChaitanya Ryali, Yuan-Ting Hu, Daniel Bolya, Chen Wei 等ICML 2023 · 被引用 388 次
- CRIS: CLIP-Driven Referring Image SegmentationZhaoqing Wang, Yu Lu, Qiang Li, Xunqiang Tao 等CVPR 2022 · 被引用 337 次
- LAVT: Language-Aware Vision Transformer for Referring Image SegmentationZhao Yang, Jiaqi Wang, Yansong Tang, Kai Chen 等CVPR 2022 · 被引用 319 次
- EfficientSAM: Leveraged Masked Image Pretraining for Efficient Segment AnythingYunyang Xiong, Bala Varadarajan, Lemeng Wu, Xiaoyu Xiang 等CVPR 2024 · 被引用 185 次
相关 Paper
- ReSAM: Refine, Requery, and Reinforce: Self-Prompting Point-Supervised Segmentation for Remote Sensing ImagesMuhammad Naseer SubhaniCVPR 2026 · 被引用 2 次
- MaskSAM: Auto-Prompt SAM with Mask Classification for Volumetric Medical Image SegmentationBin Xie, Hao Tang, Bin Duan, Dawen Cai 等ICCV 2025 · 被引用 7 次
- Prompt-Driven Referring Image Segmentation with Instance ContrastingChao Shang, Zichen Song, Heqian Qiu, Lanxiao Wang 等CVPR 2024 · 被引用 20 次
- Multi-Modal Segment Anything Model for Camouflaged Scene SegmentationGuangyu Ren, Hengyan Liu, Michalis Lazarou, Tania StathakiICCV 2025 · 被引用 2 次
- Towards Fine-Grained Interactive Segmentation in Images and VideosYuan Yao, Qiushi Yang, Miaomiao Cui, Liefeng BoICCV 2025 · 被引用 2 次
