Rotated Multi-Scale Interaction Network for Referring Remote Sensing Image Segmentation
Sihan Liu, Yiwei Ma, Xiaoqing Zhang, Haowei Wang, Jiayi Ji, Xiaoshuai Sun, Rongrong Ji
摘要
Referring Remote Sensing Image Segmentation (RRSIS) is a new challenge that combines computer vision and natural language processing. Traditional Referring Image Segmentation (RIS) approaches have been impeded by the complex spatial scales and orientations found in aerial imagery, leading to suboptimal segmentation results. To address these challenges, we introduce the Rotated Multi-Scale Interaction Network (RMSIN), an innovative approach designed for the unique demands of RRSIS. RMSIN incorporates an Intra-scale Interaction Module (IIM) to effectively address the fine-grained detail required at multiple scales and a Cross-scale Interaction Module (CIM) for integrating these details coherently across the network. Furthermore, RMSIN employs an Adaptive Rotated Convolution (ARC) to account for the diverse orientations of objects, a novel contribution that significantly enhances segmentation accuracy. To assess the efficacy of RMSIN, we have curated an expansive dataset comprising 17,402 image-caption-mask triplets, which is unparalleled in terms of scale and variety. This dataset not only presents the model with a wide range of spatial and rotational scenarios but also establishes a stringent benchmark for the RRSIS task, ensuring a rigorous evaluation of performance. Experimental evaluations demonstrate the exceptional performance of RM-SIN, surpassing existing state-of-the-art models by a significant margin. Datasets and code are available at https: //github.com/Lsan2401/RMSIN .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper21
- RemoteSAM: Towards Segment Anything for Earth ObservationLiang Yao, Fan Liu, Delong Chen, Chuanyi Zhang 等ACM MM 2025 · 被引用 28 次
- Towards Faithful Reasoning in Remote Sensing: A Perceptually-Grounded GeoSpatial Chain-of-Thought for Vision-Language ModelsJiaqi Liu, Lang Sun, Ronghao Fu, Bo YangICLR 2026 · 被引用 22 次
- RemoteReasoner: Towards Unifying Geospatial Reasoning WorkflowLiang Yao, Fan Liu, Hongbo Lu, Chuanyi Zhang 等AAAI 2026 · 被引用 16 次
- SegEarth-R2: Towards Comprehensive Language-guided Segmentation for Remote Sensing ImagesZepeng Xin, Kaiyu Li, Luodi Chen, Wanchen Li 等CVPR 2026 · 被引用 14 次
- UniGeoSeg: Towards Unified Open-World Segmentation for Geospatial ScenesShuo Ni, Di Wang, He Chen, Haonan Guo 等CVPR 2026 · 被引用 13 次
它引用的顶会 Paper27
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu 等ICCV 2021 · 被引用 31,683 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- MDETR - Modulated Detection for End-to-End Multi-Modal UnderstandingAishwarya Kamath, Mannat Singh, Yann LeCun, Gabriel Synnaeve 等ICCV 2021 · 被引用 1,114 次
- Learning High-Precision Bounding Box for Rotated Object Detection via Kullback-Leibler DivergenceXue Yang, Xiaojiang Yang, Jirui Yang, Qi Ming 等NeurIPS 2021 · 被引用 603 次
- A Fast and Accurate One-Stage Approach to Visual GroundingZhengyuan Yang, Boqing Gong, Liwei Wang, Wenbing Huang 等ICCV 2019 · 被引用 441 次
相关 Paper
- Exploring Efficient Open-Vocabulary Segmentation in the Remote SensingBingyu Li, Haocheng Dong, Da Zhang, Zhiyuan Zhao 等AAAI 2026 · 被引用 22 次
- Towards Open-Vocabulary Remote Sensing Image Semantic SegmentationChengyang Ye, Yunzhi Zhuge, Pingping ZhangAAAI 2025 · 被引用 31 次
- RIS-LAD: A Benchmark and Model for Referring Image Segmentation in Low-Altitude Drone ImageryKai Ye, YingShi Luan, Zhudi Chen, Guangyue Meng 等AAAI 2026
- Frequency Meets Semantics: Text-Visual Fusion with Directional Spectral Enhancement for Salient Object Detection in Optical Remote Sensing ImagesLamei Di, Bin Zhang, Yiming Wang, Wenxia ZhangACM MM 2025 · 被引用 2 次
- InterRVOS: Interaction-Aware Referring Video Object SegmentationWoojeong Jin, Seongchan Kim, Jaeho Lee, Seungryong KimCVPR 2026 · 被引用 6 次
