Denoise and Align: Diffusion-Driven Foreground Knowledge Prompting for Open-Vocabulary Temporal Action Detection
Sa Zhu, Wanqian Zhang, Lin Wang, Jinchao Zhang, Cong Wang, Bo Li
摘要
Open-Vocabulary Temporal Action Detection (OV-TAD) aims to localize and classify action segments of unseen categories in untrimmed videos, where effective alignment between action semantics and video representations is critical for accurate detection. However, existing methods struggle to mitigate the semantic imbalance between concise, abstract action labels and rich, complex video contents, inevitably introducing semantic noise and misleading cross-modal alignment. To address this challenge, we propose DFAlign, the first framework that leverages diffusion-based denoising to generate foreground knowledge for the guidance of action–video alignment. Following the 'conditioning, denoising and aligning' manner, we first introduce the Semantic-Unify Conditioning (SUC) module, which unifies action-shared and action-specific semantics as conditions for diffusion denoising. Then, the Background-Suppress Denoising (BSD) module generates foreground knowledge by progressively removing background redundancy from videos through denoising process. This foreground knowledge serves as effective intermediate semantic anchor between video and text representations, mitigating the semantic gap and enhancing the discriminability of action-relevant segments. Furthermore, we introduce the Foreground-Prompt Alignment (FPA) module to inject extracted foreground knowledge as prompt tokens into text representations, guiding model's attention towards action-relevant segments and enabling precise cross-modal alignment. Extensive experiments demonstrate that our method achieves state-of-the-art performance on two OV-TAD benchmarks. The code repository is provided as follows: https://github.com/Sasa77777779/DFAlign_SIGIR26.git.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper28
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- Directly Denoising Diffusion ModelsDan Zhang, Jingjing Wang, Feng LuoICML 2024 · 被引用 11,724 次
- Diffusion Self-Guidance for Controllable Image GenerationDave Epstein, Allan Jabri, Ben Poole, Alexei A. Efros 等NeurIPS 2023 · 被引用 411 次
- Relaxed Transformer Decoders for Direct Action Proposal GenerationJing Tan, Jiaqi Tang, Limin Wang, Gangshan WuICCV 2021 · 被引用 220 次
相关 Paper
- Decompose and Transfer: CoT-Prompting Enhanced Alignment for Open-Vocabulary Temporal Action DetectionSa Zhu, Wanqian Zhang, Lin Wang, Xiaohua Chen 等CVPR 2026
- KDA: Knowledge Diffusion Alignment with Enhanced Context for Video Temporal GroundingRan Ran, Jiwei Wei, Shiyuan He, Zeyu Ma 等ICCV 2025 · 被引用 4 次
- XMask3D: Cross-modal Mask Reasoning for Open Vocabulary 3D Semantic SegmentationZiyi Wang, Yanbo Wang, Xumin Yu, Jie Zhou 等NeurIPS 2024 · 被引用 7 次
- Generating Action-conditioned Prompts for Open-vocabulary Video Action RecognitionChengyou Jia, Minnan Luo, Xiaojun Chang, Zhuohang Dang 等ACM MM 2024 · 被引用 10 次
- DiffTAD: Temporal Action Detection with Proposal Denoising DiffusionSauradip Nag, Xiatian Zhu, Jiankang Deng, Yi-Zhe Song 等ICCV 2023 · 被引用 34 次
