Lune

CVPR2026Top-tier venue

TF-CADE: Foreground-Concentrated Text-Video Alignment for Zero-Shot Temporal Action Detection

Yearang Lee, Ho-Joong Kim, Seong-Whan Lee

2026Year

Abstract

Zero-Shot Temporal Action Detection (ZSTAD) aims to localize and recognize action instances from unseen action categories in untrimmed videos. Although existing methods have shown effectiveness by advancing architectural text-video alignment, they still struggle with capturing semantic distinctions between action classes, resulting in textirrelevant predictions. To address this issue, we propose a Text-Foreground Concentrated Alignment for zero-shot temporal action DEtector (TF-CADE) that explicitly aligns textual information with action-relevant foreground regions. Specifically, we introduce Action Concentrate Aggregation (ACA), which extracts action concentrate scores to aggregate temporally informative video segments into a foregroundweighted video embedding. This foreground concentrated alignment enhances the semantic consistency between text and video features and improves inter-class discriminability. In addition, a Certainty-based Confidence Re-weighting (CCR) strategy refines per-snippet confidence scores by leveraging foreground-aware similarity, effectively suppressing irrelevant action classes during inference. Extensive evaluations show that our TF-CADE not only achieves state-of-theart performance under in-distribution settings but also excels in cross-dataset generalization to unseen action classes.

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

lune papers fulltext 85758b70-ead6-4695-818f-befc8b8fd2d9

Builds on12

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines