The Devil is in the Object Boundary: Towards Annotation-free Instance Segmentation using Foundation Models
Cheng Shi, Sibei Yang
摘要
Foundation models, pre-trained on a large amount of data have demonstrated impressive zero-shot capabilities in various downstream tasks. However, in object detection and instance segmentation, two fundamental computer vision tasks heavily reliant on extensive human annotations, foundation models such as SAM and DINO struggle to achieve satisfactory performance. In this study, we reveal that the devil is in the object boundary, i.e., these foundation models fail to discern boundaries between individual objects. For the first time, we probe that CLIP, which has never accessed any instance-level annotations, can provide a highly beneficial and strong instance-level boundary prior in the clustering results of its particular intermediate layer. Following this surprising observation, we propose which ips up CL and SAM in a novel classification-first-then-discovery pipeline, enabling annotation-free, complex-scene-capable, open-vocabulary object detection and instance segmentation. Our Zip significantly boosts SAM's mask AP on COCO dataset by 12.5% and establishes state-of-the-art performance in various settings, including training-free, self-training, and label-efficient finetuning. Furthermore, annotation-free Zip even achieves comparable performance to the best-performing open-vocabulary object detecters using base annotations. Code is released at https://github.com/ChengShiest/Zip-Your-CLIP
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper8
- Vision Transformers Need More Than RegistersCheng Shi, Yizhou Yu, Sibei YangCVPR 2026 · 被引用 17 次
- Intervene-All-Paths: Unified Mitigation of LVLM Hallucinations across Alignment FormatsJiaye Qian, Ge Zheng, Yuchen Zhu, Sibei YangNeurIPS 2025 · 被引用 11 次
- Why LVLMs are More Prone to Hallucinations in Longer Responses: The Role of ContextGe Zheng, Jiaye Qian, Jiajin Tang, Sibei YangICCV 2025 · 被引用 2 次
- No More Sibling Rivalry: Debiasing Human-Object Interaction DetectionBin Yang, Yulin Zhang, Hong-Yu Zhou, Sibei YangICCV 2025
- Rethinking Query-based Transformer for Continual Image SegmentationYuchen Zhu, Cheng Shi, Dingyou Wang, Jiajin Tang 等CVPR 2025
它引用的顶会 Paper17
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
相关 Paper
- Towards Open-Vocabulary Semantic Segmentation Without Semantic LabelsHeeseong Shin, Chaehyun Kim, Sunghwan Hong, Seokju Cho 等NeurIPS 2024 · 被引用 32 次
- SegCLIP: Patch Aggregation with Learnable Centers for Open-Vocabulary Semantic SegmentationHuaishao Luo, Junwei Bao, Youzheng Wu, Xiaodong He 等ICML 2023 · 被引用 222 次
- Exploring Regional Clues in CLIP for Zero-Shot Semantic SegmentationYi Zhang, Meng-Hao Guo, Miao Wang, Shi-Min HuCVPR 2024 · 被引用 20 次
- Convolutions Die Hard: Open-Vocabulary Segmentation with Single Frozen Convolutional CLIPQihang Yu, Ju He, Xueqing Deng, Xiaohui Shen 等NeurIPS 2023 · 被引用 285 次
- CorrCLIP: Reconstructing Patch Correlations in CLIP for Open-Vocabulary Semantic SegmentationDengke Zhang, Fagui Liu, Quan TangICCV 2025 · 被引用 6 次
