MARIS: Marine Open-Vocabulary Instance Segmentation
Bingyu Li, Feiyu Wang, Da Zhang, Zhiyuan Zhao, Junyu Gao, Xuelong Li
Abstract
Most existing underwater instance segmentation approaches are constrained by close-vocabulary prediction, limiting their ability to recognize novel marine categories. To support evaluation, we introduce MARIS (Marine Open-Vocabulary Instance Segmentation), the first large-scale fine-grained benchmark for underwater Open-Vocabulary (OV) Instance segmentation (UOVIS), featuring a limited set of seen categories and diverse unseen categories. Although OV instance segmentation has shown promise on natural images, our analysis reveals that transfer to underwater scenes suffers from severe visual degradation (e.g., color attenuation) and semantic misalignment caused by lack underwater class definitions. To address these issues, we propose a unified framework with two complementary components. The Geometric Prior Enhancement Module (GPEM) leverages stable part-level and structural cues to maintain object consistency under degraded visual conditions. The Semantic Alignment Injection Mechanism (SAIM) enriches language embeddings with domainspecific priors, mitigating semantic ambiguity and improving recognition of unseen categories. Experiments show that our framework consistently outperforms existing OV baselines both In-Domain and Cross-Domain setting on MARIS, establishing a strong foundation for future underwater perception research. The code is Here 1 .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on28
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 7,873 citations
- BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and GenerationJunnan Li, Dongxu Li, Caiming Xiong, Steven C. H. HoiICML 2022 · 6,549 citations
- InstructBLIP: Towards General-purpose Vision-Language Models with Instruction TuningWenliang Dai, Junnan Li, Dongxu Li, Anthony Meng Huat Tiong et al.NeurIPS 2023 · 4,013 citations
- Depth Anything V2Lihe Yang, Bingyi Kang, Zilong Huang, Zhen Zhao et al.NeurIPS 2024 · 2,305 citations
Related papers
- Exploring the Underwater World Segmentation without Extra TrainingBingyu Li, Tao Huo, Da Zhang, Zhiyuan Zhao et al.CVPR 2026 · 18 citations
- Empowering DINO Representations for Underwater Instance Segmentation via Aligner and PrompterZhiyang Chen, Chen Zhang, Hao Fang, Runmin CongAAAI 2026 · 6 citations
- WaterMask: Instance Segmentation for Underwater ImageryShijie Lian, Hua Li, Runmin Cong, Suqi Li et al.ICCV 2023 · 72 citations
- UQ-Bench: A Benchmark for Evaluating Multimodal LLMs on Underwater Image Quality AssessmentJingchao Cao, Guo An, Feng Gao, Ke Gu et al.AAAI 2026
- Diving into Underwater: Segment Anything Model Guided Underwater Salient Instance Segmentation and A Large-scale DatasetShijie Lian, Ziyi Zhang, Hua Li, Wenjie Li et al.ICML 2024 · 52 citations
