ReAttnCLIP: Training-Free Open-Vocabulary Remote Sensing Image Segmentation via Re-defined Attention in CLIP
Xin Niu, Manqi Zhao, Dongsheng Jiang, Yingying Wu, Bing Su
Abstract
Remote sensing image segmentation is essential for applications such as natural disaster monitoring and precision agriculture. Open-vocabulary segmentation improves flexibility by removing fixed category constraints, enabling more fine-grained scene understanding. However, unlike CLIP's pretraining objective that emphasizes global image-text alignment, segmentation requires discriminative patch-level representations for accurate pixel-wise prediction. Consequently, the quality of attention maps in the final transformer layers is critical for modeling interactions among spatial regions. Existing methods often produce suboptimal representations when capturing the complex spatial structures of remote sensing imagery. To address this issue, we refine CLIP's attention mechanism through three modifications: (1) replacing patch-to-patch attention with intermediate-layer feature similarities to better preserve spatial structure; (2) leveraging intermediate-layer attention for class-to-patch alignment to reduce classification interference; and (3) disabling the [CLS] token's selfattention to mitigate bias. Experiments on multiple remote sensing benchmarks, including building and road extraction datasets, show that our method achieves state-of-theart performance among training-free approaches.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b2956c65-541a-4f99-bb7e-7e4094f2f3acBuilds on17
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao et al.ICCV 2023 · 13,211 citations
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou et al.ICCV 2021 · 8,921 citations
- Scaling Up Visual and Vision-Language Representation Learning With Noisy Text SupervisionChao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen et al.ICML 2021 · 5,401 citations
- Vision Transformers Need RegistersTimothée Darcet, Maxime Oquab, Julien Mairal, Piotr BojanowskiICLR 2024 · 769 citations
Related papers
- SegEarth-OV: Towards Training-Free Open-Vocabulary Segmentation for Remote Sensing ImagesKaiyu Li, Ruixun Liu, Xiangyong Cao, Xueru Bai et al.CVPR 2025
- CorrCLIP: Reconstructing Patch Correlations in CLIP for Open-Vocabulary Semantic SegmentationDengke Zhang, Fagui Liu, Quan TangICCV 2025 · 6 citations
- Target Refocusing via Attention Redistribution for Open-Vocabulary Semantic Segmentation: An Explainability PerspectiveJiahao Li, Yang Lu, Yachao Zhang, Yong Xie et al.AAAI 2026 · 3 citations
- OPMapper: Enhancing Open-Vocabulary Semantic Segmentation with Multi-Guidance InformationXuehui Wang, Chongjie Si, Xue Yang, Yuzhi Zhao et al.NeurIPS 2025 · 3 citations
- Unveiling the Knowledge of CLIP for Training-Free Open-Vocabulary Semantic SegmentationYajie Liu, Guodong Wang, Jinjin Zhang, Qingjie Liu et al.AAAI 2025 · 3 citations
