Distilling Self-Supervised Vision Transformers for Weakly-Supervised Few-Shot Classification & Segmentation
Dahyun Kang, Piotr Koniusz, Minsu Cho, Naila Murray
摘要
We address the task of weakly-supervised few-shot image classification and segmentation, by leveraging a Vision Transformer (ViT) pretrained with self-supervision. Our proposed method takes token representations from the selfsupervised ViT and leverages their correlations, via selfattention, to produce classification and segmentation predictions through separate task heads. Our model is able to effectively learn to perform classification and segmentation in the absence of pixel-level labels during training, using only image-level labels. To do this it uses attention maps, created from tokens generated by the selfsupervised ViT backbone, as pixel-level pseudo-labels. We also explore a practical setup with "mixed" supervision, where a small number of training images contains groundtruth pixel-level labels and the remaining images have only image-level labels. For this mixed setup, we propose to improve the pseudo-labels using a pseudo-label enhancer that was trained using the available ground-truth pixel-level labels. Experiments on Pascal-5 i and COCO-20 i demonstrate significant performance gains in a variety of supervision settings, and in particular when little-to-no pixel-level labels are available.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper12
- Adaptive FSS: A Novel Few-Shot Segmentation Framework via Prototype EnhancementJing Wang, Jiangyun Li, Chen Chen, Yisi Zhang 等AAAI 2024 · 被引用 24 次
- Active Learning for Semantic Segmentation with Multi-class Label QuerySehyun Hwang, Sohyun Lee, Hoyoung Kim, Minhyeon Oh 等NeurIPS 2023 · 被引用 22 次
- Pre-training with Random Orthogonal Projection Image ModelingMaryam Haghighat, Peyman Moghadam, Shaheer Mohamed, Piotr KoniuszICLR 2024 · 被引用 15 次
- HIDISC: A Hyperbolic Framework for Domain Generalization with Generalized Category DiscoveryVaibhav Rathore, Divyam Gupta, Biplab BanerjeeNeurIPS 2025 · 被引用 3 次
- Focus on Background: Exploring SAM's Potential in Few-shot Medical Image Segmentation with Background-centric PromptingYuntian Bo, Yazhou Zhu, Piotr Koniusz, Haofeng ZhangCVPR 2026 · 被引用 1 次
它引用的顶会 Paper34
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 被引用 24,064 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Bootstrap Your Own Latent - A New Approach to Self-Supervised LearningJean-Bastien Grill, Florian Strub, Florent Altché, Corentin Tallec 等NeurIPS 2020 · 被引用 9,171 次
- Training data-efficient image transformers & distillation through attentionHugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa 等ICML 2021 · 被引用 8,974 次
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou 等ICCV 2021 · 被引用 8,921 次
相关 Paper
- Know Your Attention Maps: Class-specific Token Masking for Weakly Supervised Semantic SegmentationJoëlle Hanna, Damian BorthICCV 2025 · 被引用 3 次
- Token Contrast for Weakly-Supervised Semantic SegmentationLixiang Ru, Heliang Zheng, Yibing Zhan, Bo DuCVPR 2023
- Semantic-Aware Superpixel for Weakly Supervised Semantic SegmentationSangtae Kim, Daeyoung Park, Byonghyo ShimAAAI 2023 · 被引用 35 次
- The Missing Point in Vision Transformers for Universal Image SegmentationSajjad Shahabodini, Mobina Mansoori, Farnoush Bayatmakou, Jamshid Abouei 等CVPR 2026 · 被引用 5 次
- MoRe: Class Patch Attention Needs Regularization for Weakly Supervised Semantic SegmentationZhiwei Yang, Yucong Meng, Kexue Fu, Shuo Wang 等AAAI 2025 · 被引用 14 次
