ReAttnCLIP: Training-Free Open-Vocabulary Remote Sensing Image Segmentation via Re-defined Attention in CLIP
Xin Niu, Manqi Zhao, Dongsheng Jiang, Yingying Wu, Bing Su
摘要
Remote sensing image segmentation is essential for applications such as natural disaster monitoring and precision agriculture. Open-vocabulary segmentation improves flexibility by removing fixed category constraints, enabling more fine-grained scene understanding. However, unlike CLIP's pretraining objective that emphasizes global image-text alignment, segmentation requires discriminative patch-level representations for accurate pixel-wise prediction. Consequently, the quality of attention maps in the final transformer layers is critical for modeling interactions among spatial regions. Existing methods often produce suboptimal representations when capturing the complex spatial structures of remote sensing imagery. To address this issue, we refine CLIP's attention mechanism through three modifications: (1) replacing patch-to-patch attention with intermediate-layer feature similarities to better preserve spatial structure; (2) leveraging intermediate-layer attention for class-to-patch alignment to reduce classification interference; and (3) disabling the [CLS] token's selfattention to mitigate bias. Experiments on multiple remote sensing benchmarks, including building and road extraction datasets, show that our method achieves state-of-theart performance among training-free approaches.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper17
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao 等ICCV 2023 · 被引用 13,211 次
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou 等ICCV 2021 · 被引用 8,921 次
- Scaling Up Visual and Vision-Language Representation Learning With Noisy Text SupervisionChao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen 等ICML 2021 · 被引用 5,401 次
- Vision Transformers Need RegistersTimothée Darcet, Maxime Oquab, Julien Mairal, Piotr BojanowskiICLR 2024 · 被引用 769 次
相关 Paper
- SegEarth-OV: Towards Training-Free Open-Vocabulary Segmentation for Remote Sensing ImagesKaiyu Li, Ruixun Liu, Xiangyong Cao, Xueru Bai 等CVPR 2025
- CorrCLIP: Reconstructing Patch Correlations in CLIP for Open-Vocabulary Semantic SegmentationDengke Zhang, Fagui Liu, Quan TangICCV 2025 · 被引用 6 次
- Target Refocusing via Attention Redistribution for Open-Vocabulary Semantic Segmentation: An Explainability PerspectiveJiahao Li, Yang Lu, Yachao Zhang, Yong Xie 等AAAI 2026 · 被引用 3 次
- OPMapper: Enhancing Open-Vocabulary Semantic Segmentation with Multi-Guidance InformationXuehui Wang, Chongjie Si, Xue Yang, Yuzhi Zhao 等NeurIPS 2025 · 被引用 3 次
- Unveiling the Knowledge of CLIP for Training-Free Open-Vocabulary Semantic SegmentationYajie Liu, Guodong Wang, Jinjin Zhang, Qingjie Liu 等AAAI 2025 · 被引用 3 次
