TF-CADE: Foreground-Concentrated Text-Video Alignment for Zero-Shot Temporal Action Detection
Yearang Lee, Ho-Joong Kim, Seong-Whan Lee
Abstract
Zero-Shot Temporal Action Detection (ZSTAD) aims to localize and recognize action instances from unseen action categories in untrimmed videos. Although existing methods have shown effectiveness by advancing architectural text-video alignment, they still struggle with capturing semantic distinctions between action classes, resulting in textirrelevant predictions. To address this issue, we propose a Text-Foreground Concentrated Alignment for zero-shot temporal action DEtector (TF-CADE) that explicitly aligns textual information with action-relevant foreground regions. Specifically, we introduce Action Concentrate Aggregation (ACA), which extracts action concentrate scores to aggregate temporally informative video segments into a foregroundweighted video embedding. This foreground concentrated alignment enhances the semantic consistency between text and video features and improves inter-class discriminability. In addition, a Certainty-based Confidence Re-weighting (CCR) strategy refines per-snippet confidence scores by leveraging foreground-aware similarity, effectively suppressing irrelevant action classes during inference. Extensive evaluations show that our TF-CADE not only achieves state-of-theart performance under in-distribution settings but also excels in cross-dataset generalization to unseen action classes.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 85758b70-ead6-4695-818f-befc8b8fd2d9Builds on12
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Scaling Up Visual and Vision-Language Representation Learning With Noisy Text SupervisionChao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen et al.ICML 2021 · 5,401 citations
- Distance-IoU Loss: Faster and Better Learning for Bounding Box RegressionZhaohui Zheng, Ping Wang, Wei Liu, Jinze Li et al.AAAI 2020 · 4,823 citations
- VideoMAE: Masked Autoencoders are Data-Efficient Learners for Self-Supervised Video Pre-TrainingZhan Tong, Yibing Song, Jue Wang, Limin WangNeurIPS 2022 · 2,336 citations
- HACS: Human Action Clips and Segments Dataset for Recognition and Temporal LocalizationHang Zhao, Antonio Torralba, Lorenzo Torresani, Zhicheng YanICCV 2019 · 298 citations
Related papers
- Text-Infused Attention and Foreground-Aware Modeling for Zero-Shot Temporal Action DetectionYearang Lee, Ho-Joong Kim, Seong-Whan LeeNeurIPS 2024 · 12 citations
- ZSTAD: Zero-Shot Temporal Activity DetectionLingling Zhang, Xiaojun Chang, Jun Liu, Minnan Luo et al.CVPR 2020
- Does Video-Text Pretraining Help Open-Vocabulary Online Action Detection?Qingsong Zhao, Yi Wang, Jilan Xu, Yinan He et al.NeurIPS 2024 · 16 citations
- Zero-Shot Object Detection by Semantics-Aware DETR with Adaptive Contrastive LossHuan Liu, Lu Zhang, Jihong Guan, Shuigeng ZhouACM MM 2023 · 6 citations
- ZBS: Zero-Shot Background Subtraction via Instance-Level Background Modeling and Foreground SelectionYongqi An, Xu Zhao, Tao Yu, Haiyun Gu et al.CVPR 2023
