Spatio-Temporal Pixel-Level Contrastive Learning-based Source-Free Domain Adaptation for Video Semantic Segmentation
Shao-Yuan Lo, Poojan Oza, Sumanth Chennupati, Alejandro Galindo, Vishal M. Patel
Abstract
Unsupervised Domain Adaptation (UDA) of semantic segmentation transfers labeled source knowledge to an unlabeled target domain by relying on accessing both the source and target data. However, the access to source data is often restricted or infeasible in real-world scenarios. Under the source data restrictive circumstances, UDA is less practical. To address this, recent works have explored solutions under the Source-Free Domain Adaptation (SFDA) setup, which aims to adapt a source-trained model to the target domain without accessing source data. Still, existing SFDA approaches use only image-level information for adaptation, making them sub-optimal in video applications. This paper studies SFDA for Video Semantic Segmentation (VSS), where temporal information is leveraged to address video adaptation. Specifically, we propose Spatio-Temporal Pixel-Level (STPL) contrastive learning, a novel method that takes full advantage of spatiotemporal information to tackle the absence of source data better. STPL explicitly learns semantic correlations among pixels in the spatio-temporal space, providing strong selfsupervision for adaptation to the unlabeled target domain. Extensive experiments show that STPL achieves state-ofthe-art performance on VSS benchmarks compared to current UDA and SFDA approaches. Code
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers9
- Doubly Contrastive Learning for Source-Free Domain Adaptive Person SearchYizhen Jia, Rong Quan, Yue Feng, Haiyan Chen et al.AAAI 2025 · 7 citations
- End-to-End Video Semantic Segmentation in Adverse Weather using Fusion Blocks and Temporal-Spatial Teacher-Student LearningXin Yang, Wending Yan, Michael Bi Mi, Yuan Yuan et al.NeurIPS 2024 · 6 citations
- Domain Adaptation for Large-Vocabulary Object DetectorsKai Jiang, Jiaxing Huang, Weiying Xie, Jie Lei et al.NeurIPS 2024 · 4 citations
- Bootstrapping Video Semantic Segmentation Model via Distillation-assisted Test-Time AdaptationJihun Kim, Hoyong Kwon, Hyeokjun Kweon, Kuk-Jin YoonCVPR 2026 · 3 citations
- MCAM: Multimodal Causal Analysis Model for Ego-Vehicle-Level Driving Video UnderstandingTongtong Cheng, Rongzhen Li, Yixin Xiong, Tao Zhang et al.ICCV 2025
Builds on23
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- Supervised Contrastive LearningPrannay Khosla, Piotr Teterwak, Chen Wang, Aaron Sarna et al.NeurIPS 2020 · 7,049 citations
- Do We Really Need to Access the Source Data? Source Hypothesis Transfer for Unsupervised Domain AdaptationJian Liang, Dapeng Hu, Jiashi FengICML 2020 · 1,624 citations
- Confidence Regularized Self-TrainingYang Zou, Zhiding Yu, Xiaofeng Liu, B. V. K. Vijaya Kumar et al.ICCV 2019 · 901 citations
- Exploring Cross-Image Pixel Contrast for Semantic SegmentationWenguan Wang, Tianfei Zhou, Fisher Yu, Jifeng Dai et al.ICCV 2021 · 568 citations
Related papers
- Source-Free Domain Adaptation for Semantic SegmentationYuang Liu, Wei Zhang, Jun WangCVPR 2021
- Spatio-temporal Contrastive Domain Adaptation for Action RecognitionXiaolin Song, Sicheng Zhao, Jingyu Yang, Huanjing Yue et al.CVPR 2021
- Source-Free Video Domain Adaptation with Spatial-Temporal-Historical Consistency LearningKai Li, Deep Patel, Erik Kruus, Martin Renqiang MinCVPR 2023
- Discovering Informative and Robust Positives for Video Domain AdaptationChang Liu, Kunpeng Li, Michael Stopa, Jun Amano et al.ICLR 2023
- CLDA: Contrastive Learning for Semi-Supervised Domain AdaptationAnkit SinghNeurIPS 2021 · 153 citations
