Diving into Underwater: Segment Anything Model Guided Underwater Salient Instance Segmentation and A Large-scale Dataset
Shijie Lian, Ziyi Zhang, Hua Li, Wenjie Li, Laurence Tianruo Yang, Sam Kwong, Runmin Cong
Abstract
With the breakthrough of large models, Segment Anything Model (SAM) and its extensions have been attempted to apply in diverse tasks of computer vision. Underwater salient instance segmentation is a foundational and vital step for various underwater vision tasks, which often suffer from low segmentation accuracy due to the complex underwater circumstances and the adaptive ability of models. Moreover, the lack of large-scale datasets with pixel-level salient instance annotations has impeded the development of machine learning techniques in this field. To address these issues, we construct the first large-scale underwater salient instance segmentation dataset (USIS10K), which contains 10,632 underwater images with pixel-level annotations in 7 categories from various underwater scenes. Then, we propose an Underwater Salient Instance Segmentation architecture based on Segment Anything Model (USIS-SAM) specifically for the underwater domain. We devise an Underwater Adaptive Visual Transformer (UA-ViT) encoder to incorporate underwater domain visual prompts into the segmentation network. We further design an out-of-the-box underwater Salient Feature Prompter Generator (SFPG) to automatically generate salient prompters instead of explicitly providing foreground points or boxes as prompts in SAM. Comprehensive experimental results show that our USIS-SAM method can achieve superior performance on USIS10K datasets compared to the state-of-the-art methods. Datasets and codes are released on https://github.com/LiamLian0727/USIS10K.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 02e9e982-2af1-4f37-9c47-a2140dcb9673Cited by top-tier papers7
- Exploring the Underwater World Segmentation without Extra TrainingBingyu Li, Tao Huo, Da Zhang, Zhiyuan Zhao et al.CVPR 2026 · 18 citations
- NAUTILUS: A Large Multimodal Model for Underwater Scene UnderstandingWei Xu, Cheng Wang, Dingkang Liang, Zongchuang Zhao et al.NeurIPS 2025 · 16 citations
- UIS-Mamba: Exploring Mamba for Underwater Instance Segmentation via Dynamic Tree Scan and Hidden State WeakenRunmin Cong, Zongji Yu, Hao Fang, Haoyan Sun et al.ACM MM 2025 · 7 citations
- Empowering DINO Representations for Underwater Instance Segmentation via Aligner and PrompterZhiyang Chen, Chen Zhang, Hao Fang, Runmin CongAAAI 2026 · 6 citations
- MARIS: Marine Open-Vocabulary Instance SegmentationBingyu Li, Feiyu Wang, Da Zhang, Zhiyuan Zhao et al.CVPR 2026
Builds on7
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao et al.ICCV 2023 · 13,211 citations
- Personalize Segment Anything Model with One ShotRenrui Zhang, Zhengkai Jiang, Ziyu Guo, Shilin Yan et al.ICLR 2024 · 333 citations
- Instances as QueriesYuxin Fang, Shusheng Yang, Xinggang Wang, Yu Li et al.ICCV 2021 · 331 citations
- Underwater Ranker: Learn Which Is Better and How to Be BetterChunle Guo, Ruiqi Wu, Xin Jin, Linghao Han et al.AAAI 2023 · 224 citations
Related papers
- BiPA: Bilevel Prompt Adaptation for Underwater Instance SegmentationLong Ma, Haoze Zheng, Yuhang Mao, Jinyuan Liu et al.CVPR 2026
- Point-SAM: Promptable 3D Segmentation Model for Point CloudsYuchen Zhou, Jiayuan Gu, Tung Yen Chiang, Fanbo Xiang et al.ICLR 2025
- Fantastic Animals and Where to Find Them: Segment Any Marine Animal with Dual SAMPingping Zhang, Tianyu Yan, Yang Liu, Huchuan LuCVPR 2024 · 32 citations
- WaterMask: Instance Segmentation for Underwater ImageryShijie Lian, Hua Li, Runmin Cong, Suqi Li et al.ICCV 2023 · 72 citations
- Towards a Comprehensive, Efficient and Promptable Anatomic Structure Segmentation Model Using 3D Whole-Body CT ScansHeng Guo, Jianfeng Zhang, Jiaxing Huang, Tony C. W. Mok et al.AAAI 2025 · 12 citations
