Fine-Grained Image-Text Correspondence with Cost Aggregation for Open-Vocabulary Part Segmentation
Jiho Choi, Seonho Lee, Minhyun Lee, Seungho Lee, Hyunjung Shim
摘要
Open-Vocabulary Part Segmentation (OVPS) is an emerging field for recognizing fine-grained parts in unseen categories. We identify two primary challenges in OVPS: (1) the difficulty in aligning part-level image-text correspondence, and (2) the lack of structural understanding in segmenting object parts. To address these issues, we propose Part-CATSeg, a novel framework that integrates object-aware part-level cost aggregation, compositional loss, and structural guidance from DINO. Our approach employs a disentangled cost aggregation strategy that handles object and part-level costs separately, enhancing the precision of partlevel segmentation. We also introduce a compositional loss to better capture part-object relationships, compensating for the limited part annotations. Additionally, structural guidance from DINO features improves boundary delineation and inter-part understanding. Extensive experiments * Equal contribution † Corresponding author on Pascal-Part-116, ADE20K-Part-234, and PartImageNet datasets demonstrate that our method significantly outperforms state-of-the-art approaches, setting a new baseline for robust generalization to unseen part categories.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- LangHOPS: Language Grounded Hierarchical Open-Vocabulary Part SegmentationYang Miao, Jan-Nico Zaech, Xi Wang, Fabien Despinoy 等NeurIPS 2025 · 被引用 3 次
- PartCo: Part-Level Correspondence Priors Enhance Category DiscoveryFernando Julio Cendra, Kai HanICML 2026 · 被引用 2 次
- Open-Vocabulary Part Segmentation via Progressive and Boundary-Aware StrategyXinlong Li, Di Lin, Shaoyiyi Gao, Jiaxin Li 等NeurIPS 2025 · 被引用 1 次
- PowerCLIP: Powerset Alignment for Contrastive Pre-TrainingMasaki Kawamura, Nakamasa Inoue, Rintaro Yanagi, Hirokatsu Kataoka 等CVPR 2026 · 被引用 1 次
- HOPS: Hierarchical Open-vocabulary Part Segmentation with Attention-Aware Filtering and Affinity-Guided EnhancementXinlong Li, Di Lin, Shaoyiyi Gao, Yaxuan Liu 等CVPR 2026
它引用的顶会 Paper37
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu 等ICCV 2021 · 被引用 31,683 次
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou 等ICCV 2021 · 被引用 8,921 次
- BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and GenerationJunnan Li, Dongxu Li, Caiming Xiong, Steven C. H. HoiICML 2022 · 被引用 6,549 次
- Scaling Up Visual and Vision-Language Representation Learning With Noisy Text SupervisionChao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen 等ICML 2021 · 被引用 5,401 次
相关 Paper
- Understanding Multi-Granularity for Open-Vocabulary Part SegmentationJiho Choi, Seonho Lee, Seungho Lee, Minhyun Lee 等NeurIPS 2024 · 被引用 7 次
- Going Denser with Open-Vocabulary Part SegmentationPeize Sun, Shoufa Chen, Chenchen Zhu, Fanyi Xiao 等ICCV 2023 · 被引用 83 次
- Knowledge-Guided Part SegmentationXuejian Gou, Fang Liu, Licheng Jiao, Shuo Li 等ICCV 2025 · 被引用 1 次
- Adaptive Part Learning for Fine-Grained Generalized Category Discovery: A Plug-and-Play EnhancementQiyuan Dai, Hanzhuo Huang, Yu Wu, Sibei YangCVPR 2025
- Towards Open-World Segmentation of PartsTai-Yu Pan, Qing Liu, Wei-Lun Chao, Brian L. PriceCVPR 2023
