RemoteSAM: Towards Segment Anything for Earth Observation
Liang Yao, Fan Liu, Delong Chen, Chuanyi Zhang, Yijun Wang, Ziyun Chen, Wei Xu, Shimin Di, Yuhui Zheng
摘要
We aim to develop a robust yet flexible visual foundation model for Earth observation. It should possess strong capabilities in recognizing and localizing diverse visual targets while providing compatibility with various input-output interfaces required across different task scenarios. Current systems cannot meet these requirements, as they typically utilize task-specific architecture trained on narrow data domains with limited semantic coverage. Our study addresses these limitations from two aspects: data and modeling. We first introduce an automatic data engine that enjoys significantly better scalability compared to previous human annotation or rule-based approaches. It has enabled us to create the largest dataset of its kind to date, comprising 270K image-text-mask triplets covering an unprecedented range of diverse semantic categories and attribute specifications. Based on this data foundation, we further propose a task unification paradigm that centers around referring expression segmentation. It effectively handles a wide range of vision-centric perception tasks, including classification, detection, segmentation, grounding, etc, using a single model without any task-specific heads. Combining these innovations on data and modeling, we present Re-moteSAM, a foundation model that establishes new SoTA on several earth observation perception benchmarks, outperforming other foundation models such as Falcon, GeoChat, and LHRS-Bot with significantly higher efficiency. Models and data are publicly available at https://github. com/1e12Leon/RemoteSAM .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper7
- RemoteReasoner: Towards Unifying Geospatial Reasoning WorkflowLiang Yao, Fan Liu, Hongbo Lu, Chuanyi Zhang 等AAAI 2026 · 被引用 16 次
- UniGeoSeg: Towards Unified Open-World Segmentation for Geospatial ScenesShuo Ni, Di Wang, He Chen, Haonan Guo 等CVPR 2026 · 被引用 13 次
- RSVG-ZeroOV: Exploring a Training-Free Framework for Zero-Shot Open-Vocabulary Visual Grounding in Remote Sensing ImagesKe Li, Di Wang, Ting Wang, Fuyu Dong 等AAAI 2026 · 被引用 7 次
- PixDLM: A Dual-Path Multimodal Language Model for UAV Reasoning SegmentationShuyan Ke, Yifan Mei, Changli Wu, Yonghan Zheng 等CVPR 2026 · 被引用 3 次
- SkySense-VITA: Towards Universal In-context Segmentation of Multi-modal Remote Sensing ImageryKang Wu, Lei Yu, Junwei Luo, Bo Dang 等CVPR 2026 · 被引用 1 次
它引用的顶会 Paper28
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu 等ICCV 2021 · 被引用 31,683 次
- MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language ModelsDeyao Zhu, Jun Chen, Xiaoqian Shen, Xiang Li 等ICLR 2024 · 被引用 3,079 次
- OFA: Unifying Architectures, Tasks, and Modalities Through a Simple Sequence-to-Sequence Learning FrameworkPeng Wang, An Yang, Rui Men, Junyang Lin 等ICML 2022 · 被引用 1,058 次
- SatMAE: Pre-training Transformers for Temporal and Multi-Spectral Satellite ImageryYezhen Cong, Samar Khanna, Chenlin Meng, Patrick Liu 等NeurIPS 2022 · 被引用 707 次
- Pix2seq: A Language Modeling Framework for Object DetectionTing Chen, Saurabh Saxena, Lala Li, David J. Fleet 等ICLR 2022 · 被引用 435 次
相关 Paper
- InstructSAM: A Training-free Framework for Instruction-Oriented Remote Sensing Object RecognitionYijie Zheng, Weijie Wu, Qingyun Li, Xuehui Wang 等NeurIPS 2025 · 被引用 12 次
- OlmoEarth: Stable Latent Image Modeling for Multimodal Earth ObservationHenry Herzog, Favyen Bastani, Yawen Zhang, Gabriel Tseng 等CVPR 2026 · 被引用 28 次
- DiffusionSat: A Generative Foundation Model for Satellite ImagerySamar Khanna, Patrick Liu, Linqi Zhou, Chenlin Meng 等ICLR 2024 · 被引用 173 次
- Unveiling Parts Beyond Objects: Towards Finer-Granularity Referring Expression SegmentationWenxuan Wang, Tongtian Yue, Yisi Zhang, Longteng Guo 等CVPR 2024 · 被引用 5 次
- Active Learning Meets Foundation Models: Fast Remote Sensing Data Annotation for Object DetectionMarvin Burges, Philipe Ambrozio Dias, Carson Woody, Sarah Walters 等ICCV 2025 · 被引用 3 次
