Scene-Aware Spatiotemporal Generalization: Towards Robust Temporal Action Detection Across Domains
Fangming Feng, Sihang Cai, Zequn Xie, Yangyang Wu, Tao Jin
Abstract
Temporal Action Detection (TAD) aims to identify specific actions in long, untrimmed videos by determining their start, end times and categories, yet existing models suffer from performance degradation under out-of-distribution scenarios due to unrealistic i.i.d. assumptions. While domain generalization (DG) offers a promising solution, image-based DG methods fail to address the unique spatiotemporal challenges in video-based TAD, including the spatiotemporal complexities and significant variations in action instance scales and densities across domains. To bridge this gap, we propose the first DG framework tailored for TAD. We propose Scene-Aware Video Segmentation, which segments videos based on semantic similarity, addressing cross-domain action instance density and scale discrepancies. Additionally, we present Temporal-Aware Normalization Perturbation to generate diverse video features while preserving temporal integrity. We establish the first DG-TAD benchmark, evaluating 11 state-of-the-art DG methods across four datasets. The experiments demonstrate that our framework consistently outperforms existing approaches, achieving superior generalization on unseen domains. The proposed modules are architecture-agnostic, offering plug-and-play compatibility for broader video understanding tasks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 911a1ca9-c12d-4e7f-b59c-3fe06c261f24Cited by top-tier papers1
Ask how each one uses itBuilds on15
- Sharpness-aware Minimization for Efficiently Improving GeneralizationPierre Foret, Ariel Kleiner, Hossein Mobahi, Behnam NeyshaburICLR 2021 · 1,861 citations
- Domain Generalization using Causal MatchingDivyat Mahajan, Shruti Tople, Amit SharmaICML 2021 · 399 citations
- Adversarial Alignment for Source Free Object DetectionQiaosong Chu, Shuyan Li, Guangyi Chen, Kai Li et al.AAAI 2023 · 62 citations
- Adversarial Bayesian Augmentation for Single-Source Domain GeneralizationSheng Cheng, Tejas Gokhale, Yezhou YangICCV 2023 · 35 citations
- Multimodal Prompt Learning with Missing Modalities for Sentiment Analysis and Emotion RecognitionZirun Guo, Tao Jin, Zhou ZhaoACL 2024 · 33 citations
Related papers
- Diversifying Spatial-Temporal Perception for Video Domain GeneralizationKun-Yu Lin, Jia-Run Du, Yipeng Gao, Jiaming Zhou et al.NeurIPS 2023 · 27 citations
- Benchmarking the Robustness of Temporal Action Detection Models Against Temporal CorruptionsRunhao Zeng, Xiaoyong Chen, Jiaming Liang, Huisi Wu et al.CVPR 2024 · 6 citations
- Action Segmentation With Joint Self-Supervised Temporal Domain AdaptationMin-Hung Chen, Baopu Li, Yingze Bao, Ghassan AlRegib et al.CVPR 2020
- Generalizable Fourier Augmentation for Unsupervised Video Object SegmentationHuihui Song, Tiankang Su, Yuhui Zheng, Kaihua Zhang et al.AAAI 2024 · 15 citations
- Tracking and Segmenting Anything in Any ModalityTianlu Zhang, Qiang Zhang, Guiguang Ding, Jungong HanAAAI 2026
