Revisiting Unsupervised Temporal Action Localization: The Primacy of High-Quality Actionness and Pseudolabels
Han Jiang, Haoyu Tang, Ming Yan, Ji Zhang, Mingzhu Xu, Yupeng Hu, Jihua Zhu, Liqiang Nie
摘要
Recently, temporal action localization (TAL) methods, especially the weakly-supervised and unsupervised ones, have become a hot research topic. Existing unsupervised methods follow an iterative ''clustering and training'' strategy with diverse model designs during training stage, while they often overlook maintaining consistency between these stages, which is crucial: more accurate clustering results can reduce the noises of pseudolabels and thus enhance model training, while more robust training can in turn enrich clustering feature representation. We identify two critical challenges in unsupervised scenarios: 1. What features should the model generate for clustering? 2. Which pseudolabeled instances from clustering should be chosen for model training? After extensive explorations, we proposed a novel yet simple framework called Consistency-Oriented Progressive high actionness Learning to address these issues. For feature generation, our framework adopts a High Actionness snippet Selection (HAS) module to generate more discriminative global video features for clustering from the enhanced actionness features obtained from a designed Inner-Outer Consistency Network (IOCNet). For pseudolabel selection, we introduces a Progressive Learning With Representative Instances (PLRI) strategy to identify the most reliable and informative instances within each cluster for model training. These three modules, HAS, IOCNet, and PLRI, synergistically improve consistency in model training and clustering performance. Extensive experiments on THUMOS'14 and ActivityNet v1.2 datasets under both unsupervised and weakly-supervised settings demonstrate that our framework achieves the state-of-the-art results.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
引用它的顶会 Paper2
- Boundary-Aware Temporal Dynamic Pseudo-Supervision Pairs Generation for Zero-Shot Natural Language Video LocalizationXiongwen Deng, Haoyu Tang, Han Jiang, Qinghai Zheng 等AAAI 2025
- Memory Matters: Boosting Training-Free Zero-Shot Temporal Action Localization with a Learnable Lookup TableHan Jiang, Haoyu Tang, Xiaoxuan Mu, Chen Li 等CVPR 2026
相关 Paper
- Similar Modality Enhancement and Action Consistency Learning for Weakly Supervised Temporal Action LocalizationMaodong Li, Chao Zheng, Jian Wang, Bing LiAAAI 2025 · 被引用 2 次
- Learning from Noisy Pseudo Labels for Semi-Supervised Temporal Action LocalizationKun Xia, Le Wang, Sanping Zhou, Gang Hua 等ICCV 2023 · 被引用 16 次
- Learning Temporal Co-Attention Models for Unsupervised Video Action LocalizationGuoqiang Gong, Xinghan Wang, Yadong Mu, Qi TianCVPR 2020
- Actionness Inconsistency-Guided Contrastive Learning for Weakly-Supervised Temporal Action LocalizationZhilin Li, Zilei Wang, Qinying LiuAAAI 2023 · 被引用 12 次
- Foreground-Action Consistency Network for Weakly Supervised Temporal Action LocalizationLinjiang Huang, Liang Wang, Hongsheng LiICCV 2021 · 被引用 91 次
