Coarse-Fine Networks for Temporal Activity Detection in Videos
Kumara Kahatapitiya, Michael S. Ryoo
Abstract
In this paper, we introduce Coarse-Fine Networks, a twostream architecture which benefits from different abstractions of temporal resolution to learn better video representations for long-term motion. Traditional Video models process inputs at one (or few) fixed temporal resolution without any dynamic frame selection. However, we argue that, processing multiple temporal resolutions of the input and doing so dynamically by learning to estimate the importance of each frame can largely improve video representations, specially in the domain of temporal activity localization. To this end, we propose (1) 'Grid Pool', a learned temporal downsampling layer to extract coarse features, and, (2) 'Multi-stage Fusion', a spatio-temporal attention mechanism to fuse a finegrained context with the coarse features. We show that our method outperforms the state-of-the-arts for action detection in public datasets including Charades with a significantly reduced compute and memory footprint. The code is available at https://github.com/kkahatapitiya/Coarse-Fine-Networks .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext adf5d9ac-d8a4-41d3-b4d5-f3753f842e97Cited by top-tier papers17
- Self-supervised Video TransformerKanchana Ranasinghe, Muzammal Naseer, Salman Khan, Fahad Shahbaz Khan et al.CVPR 2022 · 111 citations
- MS-TCT: Multi-Scale Temporal ConvTransformer for Action DetectionRui Dai, Srijan Das, Kumara Kahatapitiya, Michael S. Ryoo et al.CVPR 2022 · 93 citations
- Learning Multi-Granular Spatio-Temporal Graph Network for Skeleton-based Action RecognitionTailin Chen, Desen Zhou, Jian Wang, Shidong Wang et al.ACM MM 2021 · 79 citations
- Action Sensitivity Learning for Temporal Action LocalizationJiayi Shao, Xiaohan Wang, Ruijie Quan, Junjun Zheng et al.ICCV 2023 · 44 citations
- PointTAD: Multi-Label Temporal Action Detection with Learnable Query PointsJing Tan, Xiaotong Zhao, Xintian Shi, Bin Kang et al.NeurIPS 2022 · 41 citations
Builds on5
- SlowFast Networks for Video RecognitionChristoph Feichtenhofer, Haoqi Fan, Jitendra Malik, Kaiming HeICCV 2019 · 4,104 citations
- SCSampler: Sampling Salient Clips From Video for Efficient Action RecognitionBruno Korbar, Du Tran, Lorenzo TorresaniICCV 2019 · 257 citations
- AssembleNet: Searching for Multi-Stream Neural Connectivity in Video ArchitecturesMichael S. Ryoo, A. J. Piergiovanni, Mingxing Tan, Anelia AngelovaICLR 2020 · 109 citations
- Action Genome: Actions As Compositions of Spatio-Temporal Scene GraphsJingwei Ji, Ranjay Krishna, Li Fei-Fei, Juan Carlos NieblesCVPR 2020
- X3D: Expanding Architectures for Efficient Video RecognitionChristoph FeichtenhoferCVPR 2020
Related papers
- DeCo: Decomposition and Reconstruction for Compositional Temporal Grounding via Coarse-to-Fine Contrastive RankingLijin Yang, Quan Kong, Hsuan-Kung Yang, Wadim Kehl et al.CVPR 2023
- Modeling Multi-Label Action Dependencies for Temporal Action LocalizationPraveen Tirupattur, Kevin Duarte, Yogesh S. Rawat, Mubarak ShahCVPR 2021
- Hierarchical Self-Attention Network for Action Localization in VideosRizard Renanda Adhi Pramono, Yie-Tarng Chen, Wen-Hsien FangICCV 2019 · 41 citations
- Two-Stream Networks for Weakly-Supervised Temporal Action Localization with Semantic-Aware MechanismsYu Wang, Yadong Li, Hongbin WangCVPR 2023
- Proposal-Free Video Grounding with Contextual Pyramid NetworkKun Li, Dan Guo, Meng WangAAAI 2021 · 138 citations
