Sparse and Complete Latent Organization for Geospatial Semantic Segmentation
Fengyu Yang, Chenyang Ma
摘要
Geospatial semantic segmentation on remote sensing images suffers from large intra-class variance in both fore-ground and background classes. First, foreground objects are tiny in the remote sensing images and are represented by only a few pixels, which leads to large foreground intra-class variance and undermines the discrimination between foreground classes (issue firstly considered in this work). Second, background class contains complex context, which results in false alarms due to large background intra-class variance. To alleviate these two issues, we construct a sparse and complete latent structure via prototypes. In particular, to enhance the sparsity of the latent space, we design a prototypical contrastive learning to have prototypes of the same category clustering together and prototypes of different categories to be far away from each other. Also, we strengthen the completeness of the latent space by modeling all foreground categories and hardest (nearest) background objects. We further design a patch shuffle augmentation for remote sensing images with complicated contexts. Our augmentation encourages the semantic information of an object to be correlated only to the limited context within the patch that is specific to its category, which further reduces large intra-class variance. We conduct extensive evaluations on a large scale remote sensing dataset, showing our approach significantly outperforms state-of-the-art methods by a large margin.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper8
- Can We Leave Deepfake Data Behind in Training Deepfake Detector?Jikang Cheng, Zhiyuan Yan, Ying Zhang, Yuhao Luo 等NeurIPS 2024 · 被引用 85 次
- SpatialPIN: Enhancing Spatial Reasoning Capabilities of Vision-Language Models through Prompting and Interacting 3D PriorsChenyang Ma, Kai Lu, Ta Ying Cheng, Niki Trigoni 等NeurIPS 2024 · 被引用 82 次
- Binding Touch to Everything: Learning Unified Multimodal Tactile RepresentationsFengyu Yang, Chao Feng, Ziyang Chen, Hyoungseob Park 等CVPR 2024 · 被引用 47 次
- Generating Visual Scenes from TouchFengyu Yang, Jiacheng Zhang, Andrew OwensICCV 2023 · 被引用 39 次
- SatSynth: Augmenting Image-Mask Pairs Through Diffusion Models for Aerial Semantic SegmentationAysim Toker, Marvin Eisenberger, Daniel Cremers, Laura Leal-TaixéCVPR 2024 · 被引用 36 次
它引用的顶会 Paper14
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 被引用 24,064 次
- CutMix: Regularization Strategy to Train Strong Classifiers With Localizable FeaturesSangdoo Yun, Dongyoon Han, Sanghyuk Chun, Seong Joon Oh 等ICCV 2019 · 被引用 5,843 次
- CCNet: Criss-Cross Attention for Semantic SegmentationZilong Huang, Xinggang Wang, Lichao Huang, Chang Huang 等ICCV 2019 · 被引用 2,972 次
- Invariant Information Clustering for Unsupervised Image Classification and SegmentationXu Ji, Andrea Vedaldi, João F. HenriquesICCV 2019 · 被引用 956 次
- Gated-SCNN: Gated Shape CNNs for Semantic SegmentationTowaki Takikawa, David Acuna, Varun Jampani, Sanja FidlerICCV 2019 · 被引用 710 次
相关 Paper
- Foreground-Aware Relation Network for Geospatial Object Segmentation in High Spatial Resolution Remote Sensing ImageryZhuo Zheng, Yanfei Zhong, Junjue Wang, Ailong MaCVPR 2020
- Geography-Aware Self-Supervised LearningKumar Ayush, Burak Uzkent, Chenlin Meng, Kumar Tanmay 等ICCV 2021 · 被引用 304 次
- Continual Semantic Segmentation via Repulsion-Attraction of Sparse and Disentangled Latent RepresentationsUmberto Michieli, Pietro ZanuttighCVPR 2021
- Structure-Adaptive Multi-View Graph Clustering for Remote Sensing DataRenxiang Guan, Wenxuan Tu, Siwei Wang, Jiyuan Liu 等AAAI 2025 · 被引用 26 次
- HCSC: Hierarchical Contrastive Selective CodingYuanfan Guo, Minghao Xu, Jiawen Li, Bingbing Ni 等CVPR 2022 · 被引用 76 次
