Lune

CVPR2026顶会

ReAttnCLIP: Training-Free Open-Vocabulary Remote Sensing Image Segmentation via Re-defined Attention in CLIP

Xin Niu, Manqi Zhao, Dongsheng Jiang, Yingying Wu, Bing Su

出版方
2026年份
5被引次数

摘要

Remote sensing image segmentation is essential for applications such as natural disaster monitoring and precision agriculture. Open-vocabulary segmentation improves flexibility by removing fixed category constraints, enabling more fine-grained scene understanding. However, unlike CLIP's pretraining objective that emphasizes global image-text alignment, segmentation requires discriminative patch-level representations for accurate pixel-wise prediction. Consequently, the quality of attention maps in the final transformer layers is critical for modeling interactions among spatial regions. Existing methods often produce suboptimal representations when capturing the complex spatial structures of remote sensing imagery. To address this issue, we refine CLIP's attention mechanism through three modifications: (1) replacing patch-to-patch attention with intermediate-layer feature similarities to better preserve spatial structure; (2) leveraging intermediate-layer attention for class-to-patch alignment to reduce classification interference; and (3) disabling the [CLS] token's selfattention to mitigate bias. Experiments on multiple remote sensing benchmarks, including building and road extraction datasets, show that our method achieves state-of-theart performance among training-free approaches.

问问这篇 Paper

智能体会读完全文。

Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。

可以从这些问题问起

智能体调用

Luneget_paper_fulltext

在 Lune 里问

免费开始,无需绑卡

它引用的顶会 Paper17

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖