Relative Alignment Network for Source-Free Multimodal Video Domain Adaptation
Yi Huang, Xiaoshan Yang, Ji Zhang, Changsheng Xu
Abstract
Video domain adaptation aims to transfer knowledge from labeled source videos to unlabeled target videos. Existing video domain adaptation methods require full access to the source videos to reduce the domain gap between the source and target videos, which are impractical in real scenarios where the source videos are not available with concerns in transmission efficiency or privacy issues. To address this problem, in this paper, we propose to solve a source-free domain adaptation task for videos where only a pre-trained source model and unlabeled target videos are available for learning a multimodal video classification model. Existing source-free domain adaptation methods cannot be directly applied to this task, since videos always suffer from domain discrepancy along both the multimodal and temporal aspects, which brings difficulties in domain adaptation especially when the source data are unavailable. In this paper, we propose a Multimodal and Temporal Relative Alignment Network (MTRAN) to deal with the above challenges. To explicitly imitate the domain shifts contained in the multimodal information and the temporal dynamics of the source and target videos, we divide the target videos into two splits according to the self-entropy values of the classification results. The low-entropy videos are deemed to be source-like while the high-entropy videos are deemed to be target-like. Then, we adopt a self-entropy-guided MixUp strategy to generate synthetic samples and hypothetical samples as instance-level based on source-like and target-like videos, and push each synthetic sample to be similar with the corresponding hypothetical sample that is slightly closer to the source-like videos than the synthetic sample by multimodal and temporal relative alignment schemes. We evaluate the proposed model on four public video datasets. The results show that our model outperforms existing state-of-the-art methods.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get ddb8eadf-d36a-4553-adf6-4475c0f5bea5Cited by top-tier papers6
- Temporal Restoration and Spatial Rewiring for Source-Free Multivariate Time Series Domain AdaptationPeiliang Gong, Yucheng Wang, Min Wu, Zhenghua Chen et al.KDD 2025 · 2 citations
- Modality-Collaborative Test-Time Adaptation for Action RecognitionBaochen Xiong, Xiaoshan Yang, Yaguang Song, Yaowei Wang et al.CVPR 2024 · 1 citation
- Bridge Then Begin Anew: Generating Target-Relevant Intermediate Model for Source-Free Visual Emotion AdaptationJiankun Zhu, Sicheng Zhao, Jing Jiang, Wenbo Tang et al.AAAI 2025 · 1 citation
- Return of Frustratingly Easy Unsupervised Video Domain AdaptationPengfei Wei, Yiqun Sun, Zhiqiang Xu, Yiping Ke et al.ICML 2026
- Probability Distribution Alignment and Low-Rank Weight Decomposition for Source-Free Domain Adaptive Brain DecodingGanxi Xu, Jinyi Long, Jia ZhangAAAI 2026
Related papers
- Uncertainty-Aware Alignment Network for Cross-Domain Video-Text RetrievalXiaoshuai Hao, Wanqian ZhangNeurIPS 2023 · 26 citations
- Source-Free Video Domain Adaptation with Spatial-Temporal-Historical Consistency LearningKai Li, Deep Patel, Erik Kruus, Martin Renqiang MinCVPR 2023
- Mix-DANN and Dynamic-Modal-Distillation for Video Domain AdaptationYuehao Yin, Bin Zhu, Jingjing Chen, Lechao Cheng et al.ACM MM 2022 · 7 citations
- Hierarchical Debiasing and Noisy Correction for Cross-domain Video Tube RetrievalJingqiao Xiu, Mengze Li, Wei Ji, Jingyuan Chen et al.ACM MM 2024 · 5 citations
- Spatial-temporal Causal Inference for Partial Image-to-video AdaptationJin Chen, Xinxiao Wu, Yao Hu, Jiebo LuoAAAI 2021 · 19 citations
