Training Vision Transformers for Semi-Supervised Semantic Segmentation
Xinting Hu, Li Jiang, Bernt Schiele
Abstract
We present S4Former, a novel approach to training Vision Transformers for Semi-Supervised Semantic Segmentation (S4). At its core, S4Former employs a Vision Transformer within a classic teacher-student framework, and then leverages three novel technical ingredients: PatchShuffle as a parameter-free perturbation technique, Patch-Adaptive Self-Attention (PASA) as a fine-grainedfeature modulation method, and the innovative Negative Class Ranking (NCR) regularization loss. Based on these regu-larization modules aligned with Transformer-specific char-acteristics across the image input, feature, and output di-mensions, S4Former exploits the Transformer's ability to capture and differentiate consistent global contextual information in unlabeled images. Overall, S4 Former not only defines a new state of the art in S4 but also maintains a streamlined and scalable architecture. Being readily compatible with existing frameworks, S4 Former achieves strong improvements (up to 4.9%) on benchmarks like Pascal VOC 2012, COCO, and Cityscapes, with varying numbers of labeled data. The code is at https://github.com/JoyHuYY1412/S4Former.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 3dea856e-9c71-4883-8685-3cbe892cd2caCited by top-tier papers5
- RSVP: Reasoning Segmentation via Visual Prompting and Multi-modal Chain-of-ThoughtYi Lu, Jiawang Cao, Yongliang Wu, Bozheng Li et al.ACL 2025 · 15 citations
- Pseudo-SD: Pseudo Controlled Stable Diffusion for Semi-Supervised and Cross-Domain Semantic SegmentationDong Zhao, Qi Zang, Shuang Wang, Nicu Sebe et al.ICCV 2025 · 3 citations
- Enhancing Multimodal In-Context Learning for Image Classification through Coreset OptimizationHuiyi Chen, Jiawei Peng, Kaihua Tang, Xin Geng et al.ACM MM 2025
- Bayesian Decomposition and Semantic Completion for Few-shot Semantic SegmentationGuangchen Shi, Yirui Wu, Wei Zhu, Tao Wang et al.CVPR 2026
- Learning Beyond Vision: Vision-Language Distillation and Edge-Aware Mix Diffusion in Semi-Supervised Semantic SegmentationRui Yang, Yunfei Bai, Yuehua Liu, Xiaomao Li et al.AAAI 2026
Related papers
- SemiCVT: Semi-Supervised Convolutional Vision Transformer for Semantic SegmentationHuimin Huang, Shiao Xie, Lanfen Lin, Ruofeng Tong et al.CVPR 2023
- Semantic-Aware Superpixel for Weakly Supervised Semantic SegmentationSangtae Kim, Daeyoung Park, Byonghyo ShimAAAI 2023 · 35 citations
- MoRe: Class Patch Attention Needs Regularization for Weakly Supervised Semantic SegmentationZhiwei Yang, Yucong Meng, Kexue Fu, Shuo Wang et al.AAAI 2025 · 14 citations
- Patch-level Representation Learning for Self-supervised Vision TransformersSukmin Yun, Hankook Lee, Jaehyung Kim, Jinwoo ShinCVPR 2022 · 52 citations
- Transformer-based Open-world Instance Segmentation with Cross-task Consistency RegularizationXizhe Xue, Dongdong Yu, Lingqiao Liu, Yu Liu et al.ACM MM 2023 · 2 citations
