Lune

CVPR2026Top-tier venue

ReAttnCLIP: Training-Free Open-Vocabulary Remote Sensing Image Segmentation via Re-defined Attention in CLIP

Xin Niu, Manqi Zhao, Dongsheng Jiang, Yingying Wu, Bing Su

2026Year
5Citations

Abstract

Remote sensing image segmentation is essential for applications such as natural disaster monitoring and precision agriculture. Open-vocabulary segmentation improves flexibility by removing fixed category constraints, enabling more fine-grained scene understanding. However, unlike CLIP's pretraining objective that emphasizes global image-text alignment, segmentation requires discriminative patch-level representations for accurate pixel-wise prediction. Consequently, the quality of attention maps in the final transformer layers is critical for modeling interactions among spatial regions. Existing methods often produce suboptimal representations when capturing the complex spatial structures of remote sensing imagery. To address this issue, we refine CLIP's attention mechanism through three modifications: (1) replacing patch-to-patch attention with intermediate-layer feature similarities to better preserve spatial structure; (2) leveraging intermediate-layer attention for class-to-patch alignment to reduce classification interference; and (3) disabling the [CLS] token's selfattention to mitigate bias. Experiments on multiple remote sensing benchmarks, including building and road extraction datasets, show that our method achieves state-of-theart performance among training-free approaches.

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

lune papers fulltext b2956c65-541a-4f99-bb7e-7e4094f2f3ac

Builds on17

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines