Coarse-Fine Networks for Temporal Activity Detection in Videos
Kumara Kahatapitiya, Michael S. Ryoo
摘要
In this paper, we introduce Coarse-Fine Networks, a twostream architecture which benefits from different abstractions of temporal resolution to learn better video representations for long-term motion. Traditional Video models process inputs at one (or few) fixed temporal resolution without any dynamic frame selection. However, we argue that, processing multiple temporal resolutions of the input and doing so dynamically by learning to estimate the importance of each frame can largely improve video representations, specially in the domain of temporal activity localization. To this end, we propose (1) 'Grid Pool', a learned temporal downsampling layer to extract coarse features, and, (2) 'Multi-stage Fusion', a spatio-temporal attention mechanism to fuse a finegrained context with the coarse features. We show that our method outperforms the state-of-the-arts for action detection in public datasets including Charades with a significantly reduced compute and memory footprint. The code is available at https://github.com/kkahatapitiya/Coarse-Fine-Networks .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper17
- Self-supervised Video TransformerKanchana Ranasinghe, Muzammal Naseer, Salman Khan, Fahad Shahbaz Khan 等CVPR 2022 · 被引用 111 次
- MS-TCT: Multi-Scale Temporal ConvTransformer for Action DetectionRui Dai, Srijan Das, Kumara Kahatapitiya, Michael S. Ryoo 等CVPR 2022 · 被引用 93 次
- Learning Multi-Granular Spatio-Temporal Graph Network for Skeleton-based Action RecognitionTailin Chen, Desen Zhou, Jian Wang, Shidong Wang 等ACM MM 2021 · 被引用 79 次
- Action Sensitivity Learning for Temporal Action LocalizationJiayi Shao, Xiaohan Wang, Ruijie Quan, Junjun Zheng 等ICCV 2023 · 被引用 44 次
- PointTAD: Multi-Label Temporal Action Detection with Learnable Query PointsJing Tan, Xiaotong Zhao, Xintian Shi, Bin Kang 等NeurIPS 2022 · 被引用 41 次
它引用的顶会 Paper5
- SlowFast Networks for Video RecognitionChristoph Feichtenhofer, Haoqi Fan, Jitendra Malik, Kaiming HeICCV 2019 · 被引用 4,104 次
- SCSampler: Sampling Salient Clips From Video for Efficient Action RecognitionBruno Korbar, Du Tran, Lorenzo TorresaniICCV 2019 · 被引用 257 次
- AssembleNet: Searching for Multi-Stream Neural Connectivity in Video ArchitecturesMichael S. Ryoo, A. J. Piergiovanni, Mingxing Tan, Anelia AngelovaICLR 2020 · 被引用 109 次
- Action Genome: Actions As Compositions of Spatio-Temporal Scene GraphsJingwei Ji, Ranjay Krishna, Li Fei-Fei, Juan Carlos NieblesCVPR 2020
- X3D: Expanding Architectures for Efficient Video RecognitionChristoph FeichtenhoferCVPR 2020
相关 Paper
- DeCo: Decomposition and Reconstruction for Compositional Temporal Grounding via Coarse-to-Fine Contrastive RankingLijin Yang, Quan Kong, Hsuan-Kung Yang, Wadim Kehl 等CVPR 2023
- Modeling Multi-Label Action Dependencies for Temporal Action LocalizationPraveen Tirupattur, Kevin Duarte, Yogesh S. Rawat, Mubarak ShahCVPR 2021
- Hierarchical Self-Attention Network for Action Localization in VideosRizard Renanda Adhi Pramono, Yie-Tarng Chen, Wen-Hsien FangICCV 2019 · 被引用 41 次
- Two-Stream Networks for Weakly-Supervised Temporal Action Localization with Semantic-Aware MechanismsYu Wang, Yadong Li, Hongbin WangCVPR 2023
- Proposal-Free Video Grounding with Contextual Pyramid NetworkKun Li, Dan Guo, Meng WangAAAI 2021 · 被引用 138 次
