Learning Affinity from Attention: End-to-End Weakly-Supervised Semantic Segmentation with Transformers
Lixiang Ru, Yibing Zhan, Baosheng Yu, Bo Du
摘要
Weakly-supervised semantic segmentation (WSSS) with image-level labels is an important and challenging task. Due to the high training efficiency, end-to-end solutions for WSSS have received increasing attention from the community. However, current methods are mainly based on convolutional neural networks and fail to explore the global information properly, thus usually resulting in incomplete object regions. In this paper, to address the aforementioned problem, we introduce Transformers, which naturally integrate global information, to generate more integral initial pseudo labels for end-to-end WSSS. Motivated by the inherent consistency between the self-attention in Transformers and the semantic affinity, we propose an Affinity from Attention (AFA) module to learn semantic affinity from the multi-head self-attention (MHSA) in Transformers. The learned affinity is then leveraged to refine the initial pseudo labels for segmentation. In addition, to efficiently derive reliable affinity labels for supervising AFA and ensure the local consistency of pseudo labels, we devise a Pixel-Adaptive Refinement module that incorporates low-level image appearance information to refine the pseudo labels. We perform extensive experiments and our method achieves 66.0% and 38.9% mIoU on the PASCAL VOC 2012 and MS COCO 2014 datasets, respectively, significantly outperforming recent end-to-end methods and several multi-stage competitors. Code is available at https://github.com/rulixiang/afa.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper51
- DiffuMask: Synthesizing Images with Pixel-level Annotations for Semantic Segmentation Using Diffusion ModelsWeijia Wu, Yuzhong Zhao, Mike Zheng Shou, Hong Zhou 等ICCV 2023 · 被引用 198 次
- Self Correspondence Distillation for End-to-End Weakly-Supervised Semantic SegmentationRongtao Xu, Changwei Wang, Jiaxi Sun, Shibiao Xu 等AAAI 2023 · 被引用 82 次
- DeMT: Deformable Mixer Transformer for Multi-Task Learning of Dense PredictionYangyang Xu, Yibo Yang, Lefei ZhangAAAI 2023 · 被引用 81 次
- Uncovering Prototypical Knowledge for Weakly Open-Vocabulary Semantic SegmentationFei Zhang, Tianfei Zhou, Boyang Li, Hao He 等NeurIPS 2023 · 被引用 46 次
- Hunting Attributes: Context Prototype-Aware Learning for Weakly Supervised Semantic SegmentationFeilong Tang, Zhongxing Xu, Zhaojun Qu, Wei Feng 等CVPR 2024 · 被引用 41 次
它引用的顶会 Paper25
- SegFormer: Simple and Efficient Design for Semantic Segmentation with TransformersEnze Xie, Wenhai Wang, Zhiding Yu, Anima Anandkumar 等NeurIPS 2021 · 被引用 9,661 次
- ViViT: A Video Vision TransformerAnurag Arnab, Mostafa Dehghani, Georg Heigold, Chen Sun 等ICCV 2021 · 被引用 2,947 次
- Vision Transformers for Dense PredictionRené Ranftl, Alexey Bochkovskiy, Vladlen KoltunICCV 2021 · 被引用 2,647 次
- Per-Pixel Classification is Not All You Need for Semantic SegmentationBowen Cheng, Alexander G. Schwing, Alexander KirillovNeurIPS 2021 · 被引用 2,196 次
- TS-CAM: Token Semantic Coupled Attention Map for Weakly Supervised Object LocalizationWei Gao, Fang Wan, Xingjia Pan, Zhiliang Peng 等ICCV 2021 · 被引用 260 次
相关 Paper
- DINO is Also a Semantic Guider: Exploiting Class-aware Affinity for Weakly Supervised Semantic SegmentationYuanchen Wu, Xiaoqiang Li, Jide Li, Kequan Yang 等ACM MM 2024 · 被引用 12 次
- CIAN: Cross-Image Affinity Net for Weakly Supervised Semantic SegmentationJunsong Fan, Zhaoxiang Zhang, Tieniu Tan, Chunfeng Song 等AAAI 2020 · 被引用 230 次
- Semantic-Aware Superpixel for Weakly Supervised Semantic SegmentationSangtae Kim, Daeyoung Park, Byonghyo ShimAAAI 2023 · 被引用 35 次
- Multi-class Token Transformer for Weakly Supervised Semantic SegmentationLian Xu, Wanli Ouyang, Mohammed Bennamoun, Farid Boussaïd 等CVPR 2022 · 被引用 275 次
- Frequency-Aware Affinity for Weakly Supervised Semantic SegmentationZiqian Yang, Xianglin Qiu, Xinqiao Zhao, Xiaolei Wang 等CVPR 2026
