Plug-in Feedback Self-Adaptive Attention in CLIP for Training-Free Open-Vocabulary Segmentation
Zhixiang Chi, Yanan Wu, Li Gu, Huan Liu, Ziqiang Wang, Yang Zhang, Yang Wang, Konstantinos N. Plataniotis
Abstract
CLIP exhibits strong visual-textual alignment but struggle with open-vocabulary segmentation due to poor localization. Prior methods enhance spatial coherence by modifying intermediate attention. But, this coherence isn't consistently propagated to the final output due to subsequent operations such as projections. Additionally, intermediate attention lacks direct interaction with text representations, such semantic discrepancy limits the full potential of CLIP.
In this work, we propose a training-free, feedback-driven self-adaptive framework that adapts output-based patchlevel correspondences back to the intermediate attention.
The output predictions, being the culmination of the model's processing, encapsulate the most comprehensive visual and textual semantics about each patch. Our approach enhances semantic consistency between internal representations and final predictions by leveraging the model's outputs as a stronger spatial coherence prior. We design key modules, including attention isolation, confidence-based pruning for sparse adaptation, and adaptation ensemble, to effectively feedback the output coherence cues. Our method functions as a plug-in module, seamlessly integrating into four state-of-the-art approaches with three backbones (ViT-B, ViT-L, ViT-H). We further validate our framework across multiple attention types (Q-K, self-self, and Proxy augmented with MAE, SAM, and DINO). Our approach consistently improves their performance across eight benchmarks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext af1c50f4-b555-4948-a6f3-a43a4d24b433Cited by top-tier papers10
- Widget2Code: From Visual Widgets to UI Code via Multimodal LLMsHouston H. Zhang, Tao Zhang, Baoze Lin, Yuanqi Xue et al.CVPR 2026 · 9 citations
- PEARL: Geometry Aligns Semantics for Training-Free Open-Vocabulary Semantic SegmentationGensheng Pei, Xiruo Jiang, Xinhao Cai, Tao Chen et al.CVPR 2026 · 3 citations
- TALON: Test-time Adaptive Learning for On-the-Fly Category DiscoveryYanan Wu, Yuhan Yan, Tailai Chen, Zhixiang Chi et al.CVPR 2026 · 3 citations
- Looking Beyond the Window: Global-Local Aligned CLIP for Training-free Open-Vocabulary Semantic SegmentationByeongCheol Lee, Hyun Seok Seong, Sangeek Hyun, Gilhan Park et al.CVPR 2026 · 2 citations
- VIP: Visual-guided Prompt Evolution for Efficient Dense Vision-Language InferenceHao Zhu, Shuo Jin, Wenbin Liao, Jiayu Xiao et al.ICML 2026 · 1 citation
Builds on32
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao et al.ICCV 2023 · 13,211 citations
- Bootstrap Your Own Latent - A New Approach to Self-Supervised LearningJean-Bastien Grill, Florian Strub, Florent Altché, Corentin Tallec et al.NeurIPS 2020 · 9,171 citations
Related papers
- Unveiling the Knowledge of CLIP for Training-Free Open-Vocabulary Semantic SegmentationYajie Liu, Guodong Wang, Jinjin Zhang, Qingjie Liu et al.AAAI 2025 · 3 citations
- ResCLIP: Residual Attention for Training-free Dense Vision-language InferenceYuhang Yang, Jinhong Deng, Wen Li, Lixin DuanCVPR 2025
- CLIPeR: Hierarchically Improving Spatial Representation of CLIP for Open-Vocabulary Semantic SegmentationLin Sun, Jiale Cao, Jin Xie, Xiaoheng Jiang et al.ICCV 2025 · 5 citations
- Feature Purification Matters: Suppressing Outlier Propagation for Training-Free Open-Vocabulary Semantic SegmentationShuo Jin, Siyue Yu, Bingfeng Zhang, Mingjie Sun et al.ICCV 2025 · 3 citations
- CLIPSelf: Vision Transformer Distills Itself for Open-Vocabulary Dense PredictionSize Wu, Wenwei Zhang, Lumin Xu, Sheng Jin et al.ICLR 2024 · 129 citations
