Open-Vocabulary Segmentation with Semantic-Assisted Calibration
Yong Liu, Sule Bai, Guanbin Li, Yitong Wang, Yansong Tang
Abstract
This paper studies open-vocabulary segmentation (OVS) through calibrating in-vocabulary and domain-biased embedding space with generalized contextual prior of CLIP. As the core of open-vocabulary understanding, alignment of visual content with the semantics of unbounded text has become the bottleneck of this field. To address this challenge, recent works propose to utilize CLIP as an additional classifier and aggregate model predictions with CLIP classification results. Despite their remarkable progress, performance of OVS methods in relevant scenarios is still unsatisfactory compared with supervised counterparts. We attribute this to the in-vocabulary embedding and domainbiased CLIP prediction. To this end, we present a Semanticassisted CAlibration Network (SCAN). In SCAN, we incorporate generalized semantic prior of CLIP into proposal embedding to avoid collapsing on known categories. Besides, a contextual shift strategy is applied to mitigate the lack of global context and unnatural background noise. With above designs, SCAN achieves state-of-the-art performance on all popular open-vocabulary segmentation benchmarks. Furthermore, we also focus on the problem of existing evaluation system that ignores semantic duplication across categories, and propose a new metric called Semantic-Guided IoU (SG-IoU). Code is available here.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers22
- Towards Open-Vocabulary Remote Sensing Image Semantic SegmentationChengyang Ye, Yunzhi Zhuge, Pingping ZhangAAAI 2025 · 31 citations
- Universal Segmentation at Arbitrary Granularity with Language InstructionYong Liu, Cairong Zhang, Yitong Wang, Jiahao Wang et al.CVPR 2024 · 15 citations
- DreamLight: Towards Harmonious and Consistent Image RelightingYong Liu, Wenpeng Xiao, Qianqian Wang, Junlin Chen et al.NeurIPS 2025 · 9 citations
- KV-Edit: Training-Free Image Editing for Precise Background PreservationTianrui Zhu, Shiyi Zhang, Jiawei Shao, Yansong TangICCV 2025 · 8 citations
- Harnessing Vision Foundation Models for High-Performance, Training-Free Open Vocabulary SegmentationYuheng Shi, Minjing Dong, Chang XuICCV 2025 · 7 citations
Builds on22
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- SegFormer: Simple and Efficient Design for Semantic Segmentation with TransformersEnze Xie, Wenhai Wang, Zhiding Yu, Anima Anandkumar et al.NeurIPS 2021 · 9,661 citations
- Scaling Up Visual and Vision-Language Representation Learning With Noisy Text SupervisionChao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen et al.ICML 2021 · 5,401 citations
- SegNeXt: Rethinking Convolutional Attention Design for Semantic SegmentationMeng-Hao Guo, Cheng-Ze Lu, Qibin Hou, Zhengning Liu et al.NeurIPS 2022 · 1,385 citations
Related papers
- CLIP-Adapted Region-to-Text Learning for Generative Open-Vocabulary Semantic SegmentationJiannan Ge, Lingxi Xie, Hongtao Xie, Pandeng Li et al.ICCV 2025 · 3 citations
- S2C2Seg: Semantic-Spatial Consistency and Category Optimization for Open-Vocabulary SegmentationYuhao Qing, Yueying Wang, Chaoyang Chen, Weidong Zhang et al.CVPR 2026
- Dual Semantic Guidance for Open Vocabulary Semantic SegmentationZhengyang Wang, Tingliang Feng, Fan Lyu, Fanhua Shang et al.CVPR 2025
- Global Knowledge Calibration for Fast Open-Vocabulary SegmentationKunyang Han, Yong Liu, Jun Hao Liew, Henghui Ding et al.ICCV 2023 · 56 citations
- Seeing the Unseen: A Semantic Alignment and Context-Aware Prompt Framework for Open-Vocabulary Camouflaged Object SegmentationPeng Ren, Tian Bai, Jing Sun, Fuming SunICCV 2025 · 4 citations
