Target Refocusing via Attention Redistribution for Open-Vocabulary Semantic Segmentation: An Explainability Perspective
Jiahao Li, Yang Lu, Yachao Zhang, Yong Xie, Fangyong Wang, Yuan Xie, Yanyun Qu
Abstract
Open-vocabulary semantic segmentation (OVSS) employs pixel-level vision-language alignment to associate category-related prompts with corresponding pixels. A key challenge is enhancing the multimodal dense prediction capability, specifically this pixel-level multimodal alignment. Although existing methods achieve promising results by leveraging CLIP’s vision-language alignment, they rarely investigate the performance boundaries of CLIP for dense prediction from an interpretability mechanisms perspective. In this work, we systematically investigate CLIP's internal mechanisms and identify a critical phenomenon: analogous to human distraction, CLIP diverts significant attention resources from target regions to irrelevant tokens. Our analysis reveals that these tokens arise from dimension-specific over-activation; filtering them enhances CLIP's dense prediction performance. Consequently, we propose Refocusing CLIP (RF-CLIP), a training-free approach that emulates human distraction-refocusing behavior to redirect attention from distraction tokens back to target regions, thereby refining CLIP's multimodal alignment granularity. Our method achieves SOTA performance on eight benchmarks while maintaining high inference efficiency.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b9d90349-5437-4ba2-b1a5-13590fd6c93aCited by top-tier papers2
- MADA-Attack: Transferable Multi-modal Attention Distraction Adversarial Attack against Vision Language ModelsZhihan Qin, Jiahao Chen, Chunyi Zhou, Yuwen Pu et al.ICML 2026
- FlowDIS: Language-Guided Dichotomous Image Segmentation with Flow MatchingAndranik Sargsyan, Shant NavasardyanCVPR 2026
Builds on34
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Vision Transformers Need RegistersTimothée Darcet, Maxime Oquab, Julien Mairal, Piotr BojanowskiICLR 2024 · 769 citations
- DINO: DETR with Improved DeNoising Anchor Boxes for End-to-End Object DetectionHao Zhang, Feng Li, Shilong Liu, Lei Zhang et al.ICLR 2023 · 753 citations
- Supervision Exists Everywhere: A Data Efficient Contrastive Language-Image Pre-training ParadigmYangguang Li, Feng Liang, Lichen Zhao, Yufeng Cui et al.ICLR 2022 · 565 citations
- DenseCLIP: Language-Guided Dense Prediction with Context-Aware PromptingYongming Rao, Wenliang Zhao, Guangyi Chen, Yansong Tang et al.CVPR 2022 · 527 citations
Related papers
- S2C2Seg: Semantic-Spatial Consistency and Category Optimization for Open-Vocabulary SegmentationYuhao Qing, Yueying Wang, Chaoyang Chen, Weidong Zhang et al.CVPR 2026
- Feature Purification Matters: Suppressing Outlier Propagation for Training-Free Open-Vocabulary Semantic SegmentationShuo Jin, Siyue Yu, Bingfeng Zhang, Mingjie Sun et al.ICCV 2025 · 3 citations
- ResCLIP: Residual Attention for Training-free Dense Vision-language InferenceYuhang Yang, Jinhong Deng, Wen Li, Lixin DuanCVPR 2025
- OPMapper: Enhancing Open-Vocabulary Semantic Segmentation with Multi-Guidance InformationXuehui Wang, Chongjie Si, Xue Yang, Yuzhi Zhao et al.NeurIPS 2025 · 3 citations
- Open-Vocabulary Semantic Segmentation with Mask-adapted CLIPFeng Liang, Bichen Wu, Xiaoliang Dai, Kunpeng Li et al.CVPR 2023
