Action Segmentation With Joint Self-Supervised Temporal Domain Adaptation
Min-Hung Chen, Baopu Li, Yingze Bao, Ghassan AlRegib, Zsolt Kira
Abstract
Despite the recent progress of fully-supervised action segmentation techniques, the performance is still not fully satisfactory. One main challenge is the problem of spatiotemporal variations (e.g. different people may perform the same activity in various ways). Therefore, we exploit unlabeled videos to address this problem by reformulating the action segmentation task as a cross-domain problem with domain discrepancy caused by spatio-temporal variations. To reduce the discrepancy, we propose Self-Supervised Temporal Domain Adaptation (SSTDA), which contains two self-supervised auxiliary tasks (binary and sequential domain prediction) to jointly align cross-domain feature spaces embedded with local and global temporal dynamics, achieving better performance than other Domain Adaptation (DA) approaches. On three challenging benchmark datasets (GTEA, 50Salads, and Breakfast), SSTDA outperforms the current state-of-the-art method by large margins (e.g. for the F1@25 score, from 59.6% to 69.1% on Breakfast, from 73.4% to 81.5% on 50Salads, and from 83.6% to 89.1% on GTEA), and requires only 65% of the labeled training data for comparable performance, demonstrating the usefulness of adapting to unlabeled target videos across variations. The source code is available at https://github.com/cmhungsteve/SSTDA .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 44ef3181-145f-461e-a90d-efb7bfdcd802Cited by top-tier papers40
- Diffusion Action SegmentationDaochang Liu, Qiyue Li, Anh-Dung Dinh, Tingting Jiang et al.ICCV 2023 · 113 citations
- Learning Cross-Modal Contrastive Features for Video Domain AdaptationDonghyun Kim, Yi-Hsuan Tsai, Bingbing Zhuang, Xiang Yu et al.ICCV 2021 · 88 citations
- Refining Action Segmentation with Hierarchical Video RepresentationsHyemin Ahn, Dongheui LeeICCV 2021 · 74 citations
- Bridge-Prompt: Towards Ordinal Action Understanding in Instructional VideosMuheng Li, Lei Chen, Yueqi Duan, Zhilan Hu et al.CVPR 2022 · 70 citations
- Temporal Alignment Networks for Long-term VideoTengda Han, Weidi Xie, Andrew ZissermanCVPR 2022 · 60 citations
Builds on2
- Temporal Attentive Alignment for Large-Scale Video Domain AdaptationMin-Hung Chen, Zsolt Kira, Ghassan Alregib, Jaekwon Yoo et al.ICCV 2019 · 205 citations
- Learning Motion in Feature Space: Locally-Consistent Deformable Convolution Networks for Fine-Grained Action DetectionKhoi-Nguyen C. Mac, Dhiraj Joshi, Raymond A. Yeh, Jinjun Xiong et al.ICCV 2019 · 44 citations
Related papers
- Spatio-temporal Contrastive Domain Adaptation for Action RecognitionXiaolin Song, Sicheng Zhao, Jingyu Yang, Huanjing Yue et al.CVPR 2021
- Discovering Informative and Robust Positives for Video Domain AdaptationChang Liu, Kunpeng Li, Michael Stopa, Jun Amano et al.ICLR 2023
- Unsupervised Video Domain Adaptation for Action Recognition: A Disentanglement PerspectivePengfei Wei, Lingdong Kong, Xinghua Qu, Yi Ren et al.NeurIPS 2023 · 39 citations
- Return of Frustratingly Easy Unsupervised Video Domain AdaptationPengfei Wei, Yiqun Sun, Zhiqiang Xu, Yiping Ke et al.ICML 2026
- Self-Supervised Learning for Semi-Supervised Temporal Action ProposalXiang Wang, Shiwei Zhang, Zhiwu Qing, Yuanjie Shao et al.CVPR 2021
