HOPS: Hierarchical Open-vocabulary Part Segmentation with Attention-Aware Filtering and Affinity-Guided Enhancement
Xinlong Li, Di Lin, Shaoyiyi Gao, Yaxuan Liu, Jixian He, Jiaxin Li, Ruonan Liu, Qing Guo, Kairui Yang, Wei Feng
摘要
Open-vocabulary part segmentation (OVPS) aims to segment objects into fine-grained parts while generalizing to unseen categories. Existing VLM-based methods face two challenges: (1) object over-segmentation, caused by overly broad semantic activations, and (2) part undersegmentation, resulting from weak fine-grained perception. To address these issues, we propose HOPS, a twostage framework for hierarchical open-vocabulary part segmentation. HOPS introduces a bidirectional semantic-structural attention fusion mechanism that integrates CLIP's semantic alignment with DINO's structural perception. In the object segmentation stage, the Attention-Aware Filtering Module (AFM) refines cross-modal similarity maps via semantic-structural attention to suppress object over-segmentation. In the part segmentation stage, the Affinity-Guided Enhancement Module (AEM) iteratively propagates part responses to progressively expand activation regions, effectively mitigating part under-segmentation. Experiments on Pascal-Part-116, ADE20K-Part-234, and PartImageNet demonstrate that HOPS achieves state-ofthe-art performance with superior generalization. Our code is available at https://github.com/TJU-IDVLab/HOPS.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper28
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu 等ICCV 2021 · 被引用 31,683 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and GenerationJunnan Li, Dongxu Li, Caiming Xiong, Steven C. H. HoiICML 2022 · 被引用 6,549 次
- Align before Fuse: Vision and Language Representation Learning with Momentum DistillationJunnan Li, Ramprasaath R. Selvaraju, Akhilesh Gotmare, Shafiq R. Joty 等NeurIPS 2021 · 被引用 2,985 次
相关 Paper
- Fine-Grained Image-Text Correspondence with Cost Aggregation for Open-Vocabulary Part SegmentationJiho Choi, Seonho Lee, Minhyun Lee, Seungho Lee 等CVPR 2025
- Understanding Multi-Granularity for Open-Vocabulary Part SegmentationJiho Choi, Seonho Lee, Seungho Lee, Minhyun Lee 等NeurIPS 2024 · 被引用 7 次
- LangHOPS: Language Grounded Hierarchical Open-Vocabulary Part SegmentationYang Miao, Jan-Nico Zaech, Xi Wang, Fabien Despinoy 等NeurIPS 2025 · 被引用 3 次
- Open-Vocabulary Part Segmentation via Progressive and Boundary-Aware StrategyXinlong Li, Di Lin, Shaoyiyi Gao, Jiaxin Li 等NeurIPS 2025 · 被引用 1 次
- LoGoSeg: Integrating Local and Global Features for Open-Vocabulary Semantic SegmentationJunyang Chen, Xiangbo Lv, Zhiqiang Kou, Xingdong Sheng 等AAAI 2026
