Open-Vocabulary Semantic Segmentation via Attribute Decomposition-Aggregation
Chaofan Ma, Yuhuan Yang, Chen Ju, Fei Zhang, Ya Zhang, Yanfeng Wang
Abstract
Open-vocabulary semantic segmentation is a challenging task that requires segmenting novel object categories at inference time. Recent studies have explored vision-language pre-training to handle this task, but suffer from unrealistic assumptions in practical scenarios, i.e., low-quality textual category names. For example, this paradigm assumes that new textual categories will be accurately and completely provided, and exist in lexicons during pre-training. However, exceptions often happen when encountering ambiguity for brief or incomplete names, new words that are not present in the pre-trained lexicons, and difficult-to-describe categories for users. To address these issues, this work proposes a novel attribute decomposition-aggregation framework, AttrSeg, inspired by human cognition in understanding new concepts. Specifically, in the decomposition stage, we decouple class names into diverse attribute descriptions to complement semantic contexts from multiple perspectives. Two attribute construction strategies are designed: using large language models for common categories, and involving manually labelling for human-invented categories. In the aggregation stage, we group diverse attributes into an integrated global description, to form a discriminative classifier that distinguishes the target object from others. One hierarchical aggregation architecture is further proposed to achieve multi-level aggregations, leveraging the meticulously designed clustering module. The final results are obtained by computing the similarity between aggregated attributes and images embeddings. To evaluate the effectiveness, we annotate three types of datasets with attribute descriptions, and conduct extensive experiments and ablation studies. The results show the superior performance of attribute decomposition-aggregation.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 8f0a7924-ba34-48a3-a0ff-11364ef7ce91Cited by top-tier papers6
- Uncovering Prototypical Knowledge for Weakly Open-Vocabulary Semantic SegmentationFei Zhang, Tianfei Zhou, Boyang Li, Hao He et al.NeurIPS 2023 · 46 citations
- LLaFS: When Large Language Models Meet Few-Shot SegmentationLanyun Zhu, Tianrun Chen, Deyi Ji, Jieping Ye et al.CVPR 2024 · 39 citations
- FOLDER: Accelerating Multi-Modal Large Language Models with Enhanced PerformanceHaicheng Wang, Zhemeng Yu, Gabriele Spadaro, Chen Ju et al.ICCV 2025 · 3 citations
- Open-YOLO 3D: Towards Fast and Accurate Open-Vocabulary 3D Instance SegmentationMohamed El Amine Boudjoghra, Angela Dai, Jean Lahoud, Hisham Cholakkal et al.ICLR 2025 · 3 citations
- Advancing Myopia To Holism: Fully Contrastive Language-Image Pre-trainingHaicheng Wang, Chen Ju, Weixiong Lin, Shuai Xiao et al.CVPR 2025
Builds on24
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Scaling Up Visual and Vision-Language Representation Learning With Noisy Text SupervisionChao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen et al.ICML 2021 · 5,401 citations
- MLP-Mixer: An all-MLP Architecture for VisionIlya O. Tolstikhin, Neil Houlsby, Alexander Kolesnikov, Lucas Beyer et al.NeurIPS 2021 · 3,862 citations
Related papers
- GUIDED: Granular Understanding via Identification, Detection, and Discrimination for Fine-Grained Open-Vocabulary Object DetectionJiaming Li, Zhijia Liang, Weikai Chen, Lin Ma et al.NeurIPS 2025 · 6 citations
- Stepping Out of Similar Semantic Space for Open-Vocabulary SegmentationYong Liu, Song-Li Wu, Sule Bai, Jiahao Wang et al.ICCV 2025 · 6 citations
- Training-free Open-Vocabulary Semantic Segmentation via Diverse Prototype Construction and Sub-region MatchingXuanpu Zhao, Dianmo Sheng, Zhentao Tan, Zhiwei Zhao et al.AAAI 2025 · 2 citations
- What You Perceive Is What You Conceive: A Cognition-Inspired Framework for Open Vocabulary Image SegmentationJianghang Lin, Yue Hu, Jiangtao Shen, Yunhang Shen et al.ACM MM 2025 · 1 citation
- Towards Universal Perception through Language-Guided Open-World Object DetectionZihan Wang, Yunhang Shen, Yuan Fang, Zuwei Long et al.ACM MM 2025 · 1 citation
