Fine-Grained Image-Text Correspondence with Cost Aggregation for Open-Vocabulary Part Segmentation
Jiho Choi, Seonho Lee, Minhyun Lee, Seungho Lee, Hyunjung Shim
Abstract
Open-Vocabulary Part Segmentation (OVPS) is an emerging field for recognizing fine-grained parts in unseen categories. We identify two primary challenges in OVPS: (1) the difficulty in aligning part-level image-text correspondence, and (2) the lack of structural understanding in segmenting object parts. To address these issues, we propose Part-CATSeg, a novel framework that integrates object-aware part-level cost aggregation, compositional loss, and structural guidance from DINO. Our approach employs a disentangled cost aggregation strategy that handles object and part-level costs separately, enhancing the precision of partlevel segmentation. We also introduce a compositional loss to better capture part-object relationships, compensating for the limited part annotations. Additionally, structural guidance from DINO features improves boundary delineation and inter-part understanding. Extensive experiments * Equal contribution † Corresponding author on Pascal-Part-116, ADE20K-Part-234, and PartImageNet datasets demonstrate that our method significantly outperforms state-of-the-art approaches, setting a new baseline for robust generalization to unseen part categories.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 9150c59d-3167-490c-9404-43ce362e746dCited by top-tier papers5
- LangHOPS: Language Grounded Hierarchical Open-Vocabulary Part SegmentationYang Miao, Jan-Nico Zaech, Xi Wang, Fabien Despinoy et al.NeurIPS 2025 · 3 citations
- PartCo: Part-Level Correspondence Priors Enhance Category DiscoveryFernando Julio Cendra, Kai HanICML 2026 · 2 citations
- Open-Vocabulary Part Segmentation via Progressive and Boundary-Aware StrategyXinlong Li, Di Lin, Shaoyiyi Gao, Jiaxin Li et al.NeurIPS 2025 · 1 citation
- PowerCLIP: Powerset Alignment for Contrastive Pre-TrainingMasaki Kawamura, Nakamasa Inoue, Rintaro Yanagi, Hirokatsu Kataoka et al.CVPR 2026 · 1 citation
- HOPS: Hierarchical Open-vocabulary Part Segmentation with Attention-Aware Filtering and Affinity-Guided EnhancementXinlong Li, Di Lin, Shaoyiyi Gao, Yaxuan Liu et al.CVPR 2026
Builds on37
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou et al.ICCV 2021 · 8,921 citations
- BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and GenerationJunnan Li, Dongxu Li, Caiming Xiong, Steven C. H. HoiICML 2022 · 6,549 citations
- Scaling Up Visual and Vision-Language Representation Learning With Noisy Text SupervisionChao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen et al.ICML 2021 · 5,401 citations
Related papers
- Understanding Multi-Granularity for Open-Vocabulary Part SegmentationJiho Choi, Seonho Lee, Seungho Lee, Minhyun Lee et al.NeurIPS 2024 · 7 citations
- Going Denser with Open-Vocabulary Part SegmentationPeize Sun, Shoufa Chen, Chenchen Zhu, Fanyi Xiao et al.ICCV 2023 · 83 citations
- Knowledge-Guided Part SegmentationXuejian Gou, Fang Liu, Licheng Jiao, Shuo Li et al.ICCV 2025 · 1 citation
- Adaptive Part Learning for Fine-Grained Generalized Category Discovery: A Plug-and-Play EnhancementQiyuan Dai, Hanzhuo Huang, Yu Wu, Sibei YangCVPR 2025
- Towards Open-World Segmentation of PartsTai-Yu Pan, Qing Liu, Wei-Lun Chao, Brian L. PriceCVPR 2023
