Class Token as Proxy: Optimal Transport-Assisted Proxy Learning for Weakly Supervised Semantic Segmentation
Jian Wang, Tianhong Dai, Bingfeng Zhang, Siyue Yu, Eng Gee Lim, Jimin Xiao
摘要
Weakly Supervised Semantic Segmentation (WSSS) utilizes Class Activation Maps (CAMs) to extract spatial cues from image-level labels. However, CAMs highlight only the most discriminative foreground regions, leading to incomplete results. Recent Vision Transformer-based methods leverage class-patch attention to enhance CAMs, yet they still suffer from partial activation due to the token gap: classification-focused class tokens prioritize discriminative features, while patch tokens capture both discriminative and non-discriminative characteristics. This mismatch prevents class tokens from activating all relevant features, especially when discriminative and non-discriminative regions exhibit significant differences. To address this issue, we propose Optimal Transport-assisted Proxy Learning (OTPL), a novel framework that bridges the token gap by learning adaptive proxies. OTPL introduces two key strategies: (1) optimal transport-assisted proxy learning, which combines class tokens with their most relevant patch tokens to produce comprehensive CAMs, and (2) optimal transport-enhanced contrastive learning, aligning proxies with confident patch tokens for bounded proxy exploration. Our framework overcomes the limitation of class tokens in activating patch tokens, providing more complete and accurate CAM results. Experiments on WSSS benchmarks (PAS-CAL VOC and MS COCO) demonstrate that our method significantly improves the CAM quality and achieves stateof-the-art performances.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Leveraging Class Distributions in CLIP for Weakly Supervised Semantic SegmentationZiqian Yang, Xinqiao Zhao, Xiaolei Wang, Quan Zhang 等CVPR 2026
- Frequency-Aware Affinity for Weakly Supervised Semantic SegmentationZiqian Yang, Xianglin Qiu, Xinqiao Zhao, Xiaolei Wang 等CVPR 2026
它引用的顶会 Paper34
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Multi-class Token Transformer for Weakly Supervised Semantic SegmentationLian Xu, Wanli Ouyang, Mohammed Bennamoun, Farid Boussaïd 等CVPR 2022 · 被引用 275 次
- TS-CAM: Token Semantic Coupled Attention Map for Weakly Supervised Object LocalizationWei Gao, Fang Wan, Xingjia Pan, Zhiliang Peng 等ICCV 2021 · 被引用 260 次
- Class Re-Activation Maps for Weakly-Supervised Semantic SegmentationZhaozheng Chen, Tan Wang, Xiongwei Wu, Xian-Sheng Hua 等CVPR 2022 · 被引用 223 次
- Self-supervised Image-specific Prototype Exploration for Weakly Supervised Semantic SegmentationQi Chen, Lingxiao Yang, Jianhuang Lai, Xiaohua XieCVPR 2022 · 被引用 182 次
相关 Paper
- POT: Prototypical Optimal Transport for Weakly Supervised Semantic SegmentationJian Wang, Tianhong Dai, Bingfeng Zhang, Siyue Yu 等CVPR 2025
- Token Contrast for Weakly-Supervised Semantic SegmentationLixiang Ru, Heliang Zheng, Yibing Zhan, Bo DuCVPR 2023
- Class Tokens Infusion for Weakly Supervised Semantic SegmentationSung-Hoon Yoon, Hoyong Kwon, Hyeonseong Kim, Kuk-Jin YoonCVPR 2024 · 被引用 36 次
- MoRe: Class Patch Attention Needs Regularization for Weakly Supervised Semantic SegmentationZhiwei Yang, Yucong Meng, Kexue Fu, Shuo Wang 等AAAI 2025 · 被引用 14 次
- PSDPM: Prototype-based Secondary Discriminative Pixels Mining for Weakly Supervised Semantic SegmentationXinqiao Zhao, Ziqian Yang, Tianhong Dai, Bingfeng Zhang 等CVPR 2024 · 被引用 17 次
