WOW-Seg: A Word-free Open World Segmentation Model
Danyang Li, Tianhao Wu, Bin Lin, Zhenyuan Chen, Yang Zhang, Yuxuan Li, Ming-Ming Cheng, Xiang Li
Abstract
Open world image segmentation aims to achieve precise segmentation and semantic understanding of targets within images by addressing the infinitely open set of object categories encountered in the real world. However, traditional closed-set segmentation approaches struggle to adapt to complex open world scenarios, while foundation segmentation models such as SAM exhibit notable discrepancies between their strong segmentation capabilities and relatively weaker semantic understanding. To bridge discrepancies, we propose WOW-Seg, a Word-free Open World Segmentation model for segmenting and recognizing objects from open-set categories. Specifically, WOW-Seg introduces a novel visual prompt module, Mask2Token, which transforms image masks into visual tokens and ensures their alignment with the VLLM feature space. Moreover, We introduce the Cascade Attention Mask to decouple information across different instances. This approach mitigates inter-instance interference, leading to a significant improvement in model performance. We further construct an open world region recognition test benchmark: the Region Recognition Dataset (RR-7K). With 7,662 classes, it represents the most extensive category-rich region recognition dataset to date. WOW-Seg attains strong results on the LVIS dataset, achieving a semantic similarity of 89.7 and a semantic IoU of 82.4. This performance surpasses the previous SOTA while using only one-eighth the parameter count. These results underscore the strong open world generalization capabilities of WOW-Seg. The code and related resources are available at https://github.com/AAwcAA/WOW-Seg-Meta.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on32
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao et al.ICCV 2023 · 13,211 citations
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
- SegFormer: Simple and Efficient Design for Semantic Segmentation with TransformersEnze Xie, Wenhai Wang, Zhiding Yu, Anima Anandkumar et al.NeurIPS 2021 · 9,661 citations
- SegNeXt: Rethinking Convolutional Attention Design for Semantic SegmentationMeng-Hao Guo, Cheng-Ze Lu, Qibin Hou, Zhengning Liu et al.NeurIPS 2022 · 1,385 citations
Related papers
- Training-Free Open-Ended Object Detection and Segmentation via Attention as PromptsZhiwei Lin, Yongtao Wang, Zhi TangNeurIPS 2024 · 27 citations
- OpenWorldSAM: Extending SAM2 for Universal Image Segmentation with Language PromptsShiting Xiao, Rishabh Kabra, Yuhang Li, Donghyun Lee et al.NeurIPS 2025 · 15 citations
- WeakSAM: Segment Anything Meets Weakly-supervised Instance-level RecognitionLianghui Zhu, Junwei Zhou, Yan Liu, Xin Hao et al.ACM MM 2024 · 21 citations
- The Power of Prior: Training-Free Open-Vocabulary Semantic Segmentation with LLaVABingfeng Zhang, Siyue Yu, Hui Li, Jiahua Lin et al.CVPR 2026
- Towards Open-Vocabulary Video Instance SegmentationHaochen Wang, Xiaolong Jiang, Xu Tang, Yao Hu et al.ICCV 2023 · 56 citations
