Scene-Aware Spatiotemporal Generalization: Towards Robust Temporal Action Detection Across Domains
Fangming Feng, Sihang Cai, Zequn Xie, Yangyang Wu, Tao Jin
摘要
Temporal Action Detection (TAD) aims to identify specific actions in long, untrimmed videos by determining their start, end times and categories, yet existing models suffer from performance degradation under out-of-distribution scenarios due to unrealistic i.i.d. assumptions. While domain generalization (DG) offers a promising solution, image-based DG methods fail to address the unique spatiotemporal challenges in video-based TAD, including the spatiotemporal complexities and significant variations in action instance scales and densities across domains. To bridge this gap, we propose the first DG framework tailored for TAD. We propose Scene-Aware Video Segmentation, which segments videos based on semantic similarity, addressing cross-domain action instance density and scale discrepancies. Additionally, we present Temporal-Aware Normalization Perturbation to generate diverse video features while preserving temporal integrity. We establish the first DG-TAD benchmark, evaluating 11 state-of-the-art DG methods across four datasets. The experiments demonstrate that our framework consistently outperforms existing approaches, achieving superior generalization on unseen domains. The proposed modules are architecture-agnostic, offering plug-and-play compatibility for broader video understanding tasks.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper15
- Sharpness-aware Minimization for Efficiently Improving GeneralizationPierre Foret, Ariel Kleiner, Hossein Mobahi, Behnam NeyshaburICLR 2021 · 被引用 1,861 次
- Domain Generalization using Causal MatchingDivyat Mahajan, Shruti Tople, Amit SharmaICML 2021 · 被引用 399 次
- Adversarial Alignment for Source Free Object DetectionQiaosong Chu, Shuyan Li, Guangyi Chen, Kai Li 等AAAI 2023 · 被引用 62 次
- Adversarial Bayesian Augmentation for Single-Source Domain GeneralizationSheng Cheng, Tejas Gokhale, Yezhou YangICCV 2023 · 被引用 35 次
- Multimodal Prompt Learning with Missing Modalities for Sentiment Analysis and Emotion RecognitionZirun Guo, Tao Jin, Zhou ZhaoACL 2024 · 被引用 33 次
相关 Paper
- Diversifying Spatial-Temporal Perception for Video Domain GeneralizationKun-Yu Lin, Jia-Run Du, Yipeng Gao, Jiaming Zhou 等NeurIPS 2023 · 被引用 27 次
- Benchmarking the Robustness of Temporal Action Detection Models Against Temporal CorruptionsRunhao Zeng, Xiaoyong Chen, Jiaming Liang, Huisi Wu 等CVPR 2024 · 被引用 6 次
- Action Segmentation With Joint Self-Supervised Temporal Domain AdaptationMin-Hung Chen, Baopu Li, Yingze Bao, Ghassan AlRegib 等CVPR 2020
- Generalizable Fourier Augmentation for Unsupervised Video Object SegmentationHuihui Song, Tiankang Su, Yuhui Zheng, Kaihua Zhang 等AAAI 2024 · 被引用 15 次
- Tracking and Segmenting Anything in Any ModalityTianlu Zhang, Qiang Zhang, Guiguang Ding, Jungong HanAAAI 2026
