Open-Vocabulary Semantic Segmentation with Image Embedding Balancing
Xiangheng Shan, Dongyue Wu, Guilin Zhu, Yuanjie Shao, Nong Sang, Changxin Gao
Abstract
Open-vocabulary semantic segmentation is a challenging task, which requires the model to output semantic masks of an image beyond a close-set vocabulary. Although many efforts have been made to utilize powerful CLIP models to accomplish this task, they are still easily overfitting to training classes due to the natural gaps in semantic information between training and new classes. To overcome this challenge, we propose a novel framework for open-vocabulary semantic segmentation called EBSeg, incorpo-rating an Adaptively Balanced Decoder (AdaB Decoder) and a Semantic Structure Consistency loss (SSC Loss). The AdaB Decoder is designed to generate different image embeddings for both training and new classes. Subsequently, these two types of embeddings are adaptively balanced to fully exploit their ability to recognize training classes and generalization ability for new classes. To learn a consistent semantic structure from CLIP, the SSC Loss aligns the inter-classes affinity in the image feature space with that in the text feature space of CLIP, thereby improving the generalization ability of our model. Furthermore, we employ a frozen SAM image encoder to complement the spatial information that CLIP features lack due to the low training image resolution and image-level supervision inherent in CLIP. Extensive experiments conducted across various benchmarks demonstrate that the proposed EBSeg outperforms the state-of-the-art methods. Our code and trained models will be here: https://github.com/slonetime/EBSeg.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d0ae26a5-2b23-4953-b63a-ec956358ad66Cited by top-tier papers20
- Towards Open-Vocabulary Remote Sensing Image Semantic SegmentationChengyang Ye, Yunzhi Zhuge, Pingping ZhangAAAI 2025 · 31 citations
- Exploring Efficient Open-Vocabulary Segmentation in the Remote SensingBingyu Li, Haocheng Dong, Da Zhang, Zhiyuan Zhao et al.AAAI 2026 · 22 citations
- PCA-Seg: Revisiting Cost Aggregation for Open-Vocabulary Semantic and Part SegmentationJianjian Yin, Tao Chen, Yi Chen, Gensheng Pei et al.CVPR 2026 · 6 citations
- MM-OVSeg: Multimodal Optical-SAR Fusion for Open-Vocabulary Segmentation in Remote SensingYimin Wei, Aoran Xiao, Hongruixuan Chen, Junshi Xia et al.CVPR 2026 · 6 citations
- Leveraging Depth and Language for Open-Vocabulary Domain-Generalized Semantic SegmentationSiyu Chen, Ting Han, Chengzheng Fu, Changshe Zhang et al.NeurIPS 2025 · 4 citations
Builds on18
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao et al.ICCV 2023 · 13,211 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- SegFormer: Simple and Efficient Design for Semantic Segmentation with TransformersEnze Xie, Wenhai Wang, Zhiding Yu, Anima Anandkumar et al.NeurIPS 2021 · 9,661 citations
Related papers
- Mask-Adapter: The Devil is in the Masks for Open-Vocabulary SegmentationYongkang Li, Tianheng Cheng, Bin Feng, Wenyu Liu et al.CVPR 2025
- Convolutions Die Hard: Open-Vocabulary Segmentation with Single Frozen Convolutional CLIPQihang Yu, Ju He, Xueqing Deng, Xiaohui Shen et al.NeurIPS 2023 · 285 citations
- CLIP-Adapted Region-to-Text Learning for Generative Open-Vocabulary Semantic SegmentationJiannan Ge, Lingxi Xie, Hongtao Xie, Pandeng Li et al.ICCV 2025 · 3 citations
- Seeing the Unseen: A Semantic Alignment and Context-Aware Prompt Framework for Open-Vocabulary Camouflaged Object SegmentationPeng Ren, Tian Bai, Jing Sun, Fuming SunICCV 2025 · 4 citations
- Side Adapter Network for Open-Vocabulary Semantic SegmentationMengde Xu, Zheng Zhang, Fangyun Wei, Han Hu et al.CVPR 2023
