Text-Infused Attention and Foreground-Aware Modeling for Zero-Shot Temporal Action Detection
Yearang Lee, Ho-Joong Kim, Seong-Whan Lee
摘要
Zero-Shot Temporal Action Detection (ZSTAD) aims to classify and localize action segments in untrimmed videos for unseen action categories. Most existing ZSTAD methods utilize a foreground-based approach, limiting the integration of text and visual features due to their reliance on pre-extracted proposals. In this paper, we introduce a cross-modal ZSTAD baseline with mutual cross-attention, integrating both text and visual information throughout the detection process. Our simple approach results in superior performance compared to previous methods. Despite this improvement, we further identify a common-action bias issue that the cross-modal baseline over-focus on common sub-actions due to a lack of ability to discriminate text-related visual parts. To address this issue, we propose Text-infused attention and Foreground-aware Action Detection (Ti-FAD), which enhances the ability to focus on text-related sub-actions and distinguish relevant action segments from the background. Our extensive experiments demonstrate that Ti-FAD outperforms the state-of-the-art methods on ZSTAD benchmarks by a large margin: 41.2% (+ 11.0%) on THUMOS14 and 32.0% (+ 5.4%) on ActivityNet v1.3. Code is available at: https://github.com/YearangLee/Ti-FAD.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Denoise and Align: Diffusion-Driven Foreground Knowledge Prompting for Open-Vocabulary Temporal Action DetectionSa Zhu, Wanqian Zhang, Lin Wang, Jinchao Zhang 等SIGIR 2026
- TF-CADE: Foreground-Concentrated Text-Video Alignment for Zero-Shot Temporal Action DetectionYearang Lee, Ho-Joong Kim, Seong-Whan LeeCVPR 2026
- DiGIT: Multi-Dilated Gated Encoder and Central-Adjacent Region Integrated Decoder for Temporal Action Detection TransformerHo-Joong Kim, Yearang Lee, Jung-Ho Hong, Seong-Whan LeeCVPR 2025
它引用的顶会 Paper14
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Scaling Up Visual and Vision-Language Representation Learning With Noisy Text SupervisionChao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen 等ICML 2021 · 被引用 5,401 次
- Distance-IoU Loss: Faster and Better Learning for Bounding Box RegressionZhaohui Zheng, Ping Wang, Wei Liu, Jinze Li 等AAAI 2020 · 被引用 4,823 次
- Align before Fuse: Vision and Language Representation Learning with Momentum DistillationJunnan Li, Ramprasaath R. Selvaraju, Akhilesh Gotmare, Shafiq R. Joty 等NeurIPS 2021 · 被引用 2,985 次
- AdaptFormer: Adapting Vision Transformers for Scalable Visual RecognitionShoufa Chen, Chongjian Ge, Zhan Tong, Jiangliu Wang 等NeurIPS 2022 · 被引用 1,291 次
相关 Paper
- ZSTAD: Zero-Shot Temporal Activity DetectionLingling Zhang, Xiaojun Chang, Jun Liu, Minnan Luo 等CVPR 2020
- Does Video-Text Pretraining Help Open-Vocabulary Online Action Detection?Qingsong Zhao, Yi Wang, Jilan Xu, Yinan He 等NeurIPS 2024 · 被引用 16 次
- Test-Time Zero-Shot Temporal Action LocalizationBenedetta Liberatori, Alessandro Conti, Paolo Rota, Yiming Wang 等CVPR 2024
- Building a Multi-modal Spatiotemporal Expert for Zero-shot Action Recognition with CLIPYating Yu, Congqi Cao, Yueran Zhang, Qinyi Lv 等AAAI 2025 · 被引用 12 次
- Crossmodal Representation Learning for Zero-shot Action RecognitionChung-Ching Lin, Kevin Lin, Lijuan Wang, Zicheng Liu 等CVPR 2022 · 被引用 39 次
