Grounding Everything: Emerging Localization Properties in Vision-Language Transformers
Walid Bousselham, Felix Petersen, Vittorio Ferrari, Hilde Kuehne
摘要
Vision-language foundation models have shown remarkable performance in various zero-shot settings such as image retrieval, classification, or captioning. But so far, those models seem to fall behind when it comes to zero-shot localization of referential expressions and objects in images. In this paper, we show that pretrained vision-language (VL) models allow for zero-shot open-vocabulary object localization without any fine-tuning. To leverage those capabilities, we propose a Grounding Everything Module (GEM) that generalizes the idea of value-value attention introduced by CLIPSurgery [14] to a self-self attention path. We show that the concept of self-self attention corresponds to clustering, thus enforcing groups of tokens arising from the same object to be similar while preserving the alignment with the language space. To further guide the group formation, we propose a set of regularizations that allows the model to better generalize across datasets and backbones. We evaluate the proposed GEM framework on various benchmark tasks and datasets for semantic segmentation. It shows that GEM not only outperforms other training-free open-vocabulary localization methods, but also achieves state-of-the-art results on the recently proposed OpenImagesV7 large-scale segmentation benchmark.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper45
- Talking to DINO: Bridging Self-Supervised Vision Backbones with Language for Open-Vocabulary SegmentationLuca Barsellotti, Lorenzo Bianchi, Nicola Messina, Fabio Carrara 等ICCV 2025 · 被引用 58 次
- LeGrad: An Explainability Method for Vision Transformers via Feature Formation SensitivityWalid Bousselham, Angie W. Boggust, Sofian Chaybouti, Hendrik Strobelt 等ICCV 2025 · 被引用 47 次
- RemoteSAM: Towards Segment Anything for Earth ObservationLiang Yao, Fan Liu, Delong Chen, Chuanyi Zhang 等ACM MM 2025 · 被引用 28 次
- Vision Transformers with Self-Distilled RegistersZipeng Yan, Yinjie Chen, Chong Zhou, Bo Dai 等NeurIPS 2025 · 被引用 17 次
- B-cosification: Transforming Deep Neural Networks to be Inherently InterpretableShreyash Arya, Sukrut Rao, Moritz Böhle, Bernt SchieleNeurIPS 2024 · 被引用 14 次
它引用的顶会 Paper17
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao 等ICCV 2023 · 被引用 13,211 次
- BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and GenerationJunnan Li, Dongxu Li, Caiming Xiong, Steven C. H. HoiICML 2022 · 被引用 6,549 次
- Scaling Up Visual and Vision-Language Representation Learning With Noisy Text SupervisionChao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen 等ICML 2021 · 被引用 5,401 次
- Object-Centric Learning with Slot AttentionFrancesco Locatello, Dirk Weissenborn, Thomas Unterthiner, Aravindh Mahendran 等NeurIPS 2020 · 被引用 1,275 次
相关 Paper
- VideoGEM: Training-free Action Grounding in VideosFelix Vogel, Walid Bousselham, Anna Kukleva, Nina Shvetsova 等CVPR 2025
- RSVG-ZeroOV: Exploring a Training-Free Framework for Zero-Shot Open-Vocabulary Visual Grounding in Remote Sensing ImagesKe Li, Di Wang, Ting Wang, Fuyu Dong 等AAAI 2026 · 被引用 7 次
- Exploring Open-Vocabulary Semantic Segmentation from CLIP Vision Encoder Distillation OnlyJun Chen, Deyao Zhu, Guocheng Qian, Bernard Ghanem 等ICCV 2023 · 被引用 60 次
- Unbiased Region-Language Alignment for Open-Vocabulary Dense PredictionYunheng Li, Yuxuan Li, Quan-Sheng Zeng, Wenhai Wang 等ICCV 2025 · 被引用 3 次
- Zero-guidance Segmentation Using Zero Segment LabelsPitchaporn Rewatbowornwong, Nattanat Chatthee, Ekapol Chuangsuwanich, Supasorn SuwajanakornICCV 2023 · 被引用 21 次
