Open-Vocabulary Semantic Segmentation with Decoupled One-Pass Network
Cong Han, Yujie Zhong, Dengjie Li, Kai Han, Lin Ma
Abstract
Recently, the open-vocabulary semantic segmentation problem has attracted increasing attention and the best performing methods are based on two-stream networks: one stream for proposal mask generation and the other for segment classification using a pre-trained visual-language model. However, existing two-stream methods require passing a great number of (up to a hundred) image crops into the visual-language model, which is highly inefficient. To address the problem, we propose a network that only needs a single pass through the visual-language model for each input image. Specifically, we first propose a novel network adaptation approach, termed patch severance, to restrict the harmful interference between the patch embeddings in the pre-trained visual encoder. We then propose classification anchor learning to encourage the network to spatially focus on more discriminative features for classification. Extensive experiments demonstrate that the proposed method achieves outstanding performance, surpassing state-of-the-art methods while being 4 to 7 times faster at inference. Code: https://github.com/CongHan0808/DeOP.git
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers22
- Talking to DINO: Bridging Self-Supervised Vision Backbones with Language for Open-Vocabulary SegmentationLuca Barsellotti, Lorenzo Bianchi, Nicola Messina, Fabio Carrara et al.ICCV 2025 · 58 citations
- Cascade-CLIP: Cascaded Vision-Language Embeddings Alignment for Zero-Shot Semantic SegmentationYunheng Li, Zhong-Yu Li, Quan-Sheng Zeng, Qibin Hou et al.ICML 2024 · 27 citations
- OpenVIS: Open-vocabulary Video Instance SegmentationPinxue Guo, Hao Huang, Peiyang He, Xuefeng Liu et al.AAAI 2025 · 26 citations
- Exploring Regional Clues in CLIP for Zero-Shot Semantic SegmentationYi Zhang, Meng-Hao Guo, Miao Wang, Shi-Min HuCVPR 2024 · 20 citations
- Image-to-Image Matching via Foundation Models: A New Perspective for Open-Vocabulary Semantic SegmentationYuan Wang, Rui Sun, Naisong Luo, Yuwen Pan et al.CVPR 2024 · 13 citations
Builds on15
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Scaling Up Visual and Vision-Language Representation Learning With Noisy Text SupervisionChao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen et al.ICML 2021 · 5,401 citations
- Per-Pixel Classification is Not All You Need for Semantic SegmentationBowen Cheng, Alexander G. Schwing, Alexander KirillovNeurIPS 2021 · 2,196 citations
- TOOD: Task-aligned One-stage Object DetectionChengjian Feng, Yujie Zhong, Yu Gao, Matthew R. Scott et al.ICCV 2021 · 1,191 citations
Related papers
- Side Adapter Network for Open-Vocabulary Semantic SegmentationMengde Xu, Zheng Zhang, Fangyun Wei, Han Hu et al.CVPR 2023
- SED: A Simple Encoder-Decoder for Open-Vocabulary Semantic SegmentationBin Xie, Jiale Cao, Jin Xie, Fahad Shahbaz Khan et al.CVPR 2024 · 57 citations
- CLIP-Adapted Region-to-Text Learning for Generative Open-Vocabulary Semantic SegmentationJiannan Ge, Lingxi Xie, Hongtao Xie, Pandeng Li et al.ICCV 2025 · 3 citations
- Open-vocabulary Panoptic Segmentation with Embedding ModulationXi Chen, Shuang Li, Ser-Nam Lim, Antonio Torralba et al.ICCV 2023 · 42 citations
- Mask-Adapter: The Devil is in the Masks for Open-Vocabulary SegmentationYongkang Li, Tianheng Cheng, Bin Feng, Wenyu Liu et al.CVPR 2025
