Temporal Query Networks for Fine-Grained Video Understanding
Chuhan Zhang, Ankush Gupta, Andrew Zisserman
Abstract
Our objective in this work is fine-grained classification of actions in untrimmed videos, where the actions may be temporally extended or may span only a few frames of the video. We cast this into a query-response mechanism, where each query addresses a particular question, and has its own response label set. We make the following four contributions: (i) We propose a new model-a Temporal Query Network-which enables the query-response functionality, and a structural understanding of fine-grained actions. It attends to relevant segments for each query with a temporal attention mechanism, and can be trained using only the labels for each query. (ii) We propose a new way-stochastic feature bank update-to train a network on videos of various lengths with the dense sampling required to respond to fine-grained queries. (iii) we compare the TQN to other architectures and text supervision methods, and analyze their pros and cons. Finally, (iv) we evaluate the method extensively on the FineGym and Diving48 benchmarks for fine-grained action classification and surpass the
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 5b6f907e-cfc4-4bcb-a00d-26edb6df79d9Cited by top-tier papers36
- Detecting Moments and Highlights in Videos via Natural Language QueriesJie Lei, Tamara L. Berg, Mohit BansalNeurIPS 2021 · 425 citations
- BEVT: BERT Pretraining of Video TransformersRui Wang, Dongdong Chen, Zuxuan Wu, Yinpeng Chen et al.CVPR 2022 · 200 citations
- X-Pool: Cross-Modal Language-Video Attention for Text-Video RetrievalSatya Krishna Gorti, Noël Vouitsis, Junwei Ma, Keyvan Golestan et al.CVPR 2022 · 190 citations
- MS-TCT: Multi-Scale Temporal ConvTransformer for Action DetectionRui Dai, Srijan Das, Kumara Kahatapitiya, Michael S. Ryoo et al.CVPR 2022 · 93 citations
- SkeletonMAE: Graph-based Masked Autoencoder for Skeleton Sequence Pre-trainingHong Yan, Yang Liu, Yushen Wei, Zhen Li et al.ICCV 2023 · 77 citations
Builds on18
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- SlowFast Networks for Video RecognitionChristoph Feichtenhofer, Haoqi Fan, Jitendra Malik, Kaiming HeICCV 2019 · 4,104 citations
- TSM: Temporal Shift Module for Efficient Video UnderstandingJi Lin, Chuang Gan, Song HanICCV 2019 · 2,049 citations
- Generative Pretraining From PixelsMark Chen, Alec Radford, Rewon Child, Jeffrey Wu et al.ICML 2020 · 1,773 citations
- HowTo100M: Learning a Text-Video Embedding by Watching Hundred Million Narrated Video ClipsAntoine Miech, Dimitri Zhukov, Jean-Baptiste Alayrac, Makarand Tapaswi et al.ICCV 2019 · 1,437 citations
Related papers
- Fine-grained Action Recognition with Robust Motion Representation Decoupling and ConcentrationBaoli Sun, Xinchen Ye, Tiantian Yan, Zhihui Wang et al.ACM MM 2022 · 11 citations
- FineDiving: A Fine-grained Dataset for Procedure-aware Action Quality AssessmentJinglin Xu, Yongming Rao, Xumin Yu, Guangyi Chen et al.CVPR 2022 · 118 citations
- Compact Bilinear Augmented Query Structured Attention for Sport Highlights ClassificationYanbin Hao, Hao Zhang, Chong-Wah Ngo, Qiang Liu et al.ACM MM 2020 · 20 citations
- Coarse-Fine Networks for Temporal Activity Detection in VideosKumara Kahatapitiya, Michael S. RyooCVPR 2021
- Exploring Coarse-to-Fine Action Token Localization and Interaction for Fine-grained Video Action RecognitionBaoli Sun, Xinchen Ye, Zhihui Wang, Haojie Li et al.ACM MM 2023 · 7 citations
