Diversifying Spatial-Temporal Perception for Video Domain Generalization
Kun-Yu Lin, Jia-Run Du, Yipeng Gao, Jiaming Zhou, Wei-Shi Zheng
Abstract
Video domain generalization aims to learn generalizable video classification models for unseen target domains by training in a source domain. A critical challenge of video domain generalization is to defend against the heavy reliance on domain-specific cues extracted from the source domain when recognizing target videos. To this end, we propose to perceive diverse spatial-temporal cues in videos, aiming to discover potential domain-invariant cues in addition to domain-specific cues. We contribute a novel model named Spatial-Temporal Diversification Network (STDN), which improves the diversity from both space and time dimensions of video data. First, our STDN proposes to discover various types of spatial cues within individual frames by spatial grouping. Then, our STDN proposes to explicitly model spatial-temporal dependencies between video contents at multiple space-time scales by spatial-temporal relation modeling. Extensive experiments on three benchmarks of different types demonstrate the effectiveness and versatility of our approach.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 220b066c-bd56-41b5-84d4-0130a660d1edCited by top-tier papers9
- Less Static, More Private: Towards Transferable Privacy-Preserving Action Recognition by Generative Decoupled LearningZhi-Wei Xia, Kun-Yu Lin, Yuan-Ming Li, Wei-Jin Huang et al.ICCV 2025 · 2 citations
- Generalising Traffic Forecasting to Regions Without Traffic ObservationsXinyu Su, Majid Sarvi, Feng Liu, Egemen Tanin et al.AAAI 2026 · 1 citation
- Punching Bag vs. Punching Person: Motion Transferability in VideosRaiyaan Abdullah, Jared Claypoole, Michael Cogswell, Ajay Divakaran et al.ICCV 2025 · 1 citation
- Ranking Distillation for Open-Ended Video Question Answering with Insufficient LabelsTianming Liang, Chaolei Tan, Beihao Xia, Wei-Shi Zheng et al.CVPR 2024
- Return of Frustratingly Easy Unsupervised Video Domain AdaptationPengfei Wei, Yiqun Sun, Zhiqiang Xu, Yiping Ke et al.ICML 2026
Builds on47
- ViViT: A Video Vision TransformerAnurag Arnab, Mostafa Dehghani, Georg Heigold, Chen Sun et al.ICCV 2021 · 2,947 citations
- Is Space-Time Attention All You Need for Video Understanding?Gedas Bertasius, Heng Wang, Lorenzo TorresaniICML 2021 · 2,927 citations
- TSM: Temporal Shift Module for Efficient Video UnderstandingJi Lin, Chuang Gan, Song HanICCV 2019 · 2,049 citations
- Video Swin TransformerZe Liu, Jia Ning, Yue Cao, Yixuan Wei et al.CVPR 2022 · 1,847 citations
- In Search of Lost Domain GeneralizationIshaan Gulrajani, David Lopez-PazICLR 2021 · 1,416 citations
Related papers
- Scene-Aware Spatiotemporal Generalization: Towards Robust Temporal Action Detection Across DomainsFangming Feng, Sihang Cai, Zequn Xie, Yangyang Wu et al.AAAI 2026 · 1 citation
- Revisiting Temporal Modeling for CLIP-Based Image-to-Video Knowledge TransferringRuyang Liu, Jingjia Huang, Ge Li, Jiashi Feng et al.CVPR 2023
- STDiff: Spatio-Temporal Diffusion for Continuous Stochastic Video PredictionXi Ye, Guillaume-Alexandre BilodeauAAAI 2024 · 20 citations
- Unsupervised Video Domain Adaptation for Action Recognition: A Disentanglement PerspectivePengfei Wei, Lingdong Kong, Xinghua Qu, Yi Ren et al.NeurIPS 2023 · 39 citations
- Domain Adaptive Video Segmentation via Temporal Consistency RegularizationDayan Guan, Jiaxing Huang, Aoran Xiao, Shijian LuICCV 2021 · 44 citations
