Representing Videos As Discriminative Sub-Graphs for Action Recognition
Dong Li, Zhaofan Qiu, Yingwei Pan, Ting Yao, Houqiang Li, Tao Mei
摘要
Human actions are typically of combinatorial structures or patterns, i.e., subjects, objects, plus spatio-temporal interactions in between. Discovering such structures is therefore a rewarding way to reason about the dynamics of interactions and recognize the actions. In this paper, we introduce a new design of sub-graphs to represent and encode the discriminative patterns of each action in the videos. Specifically, we present MUlti-scale Sub-graph LEarning (MUSLE) framework that novelly builds space-time graphs and clusters the graphs into compact sub-graphs on each scale with respect to the number of nodes. Technically, MUSLE produces 3D bounding boxes, i.e., tubelets, in each video clip, as graph nodes and takes dense connectivity as graph edges between tubelets. For each action category, we execute online clustering to decompose the graph into sub-graphs on each scale through learning Gaussian Mixture Layer and select the discriminative sub-graphs as action prototypes for recognition. Extensive experiments are conducted on both Something-Something V1 & V2 and Kinetics-400 datasets, and superior results are reported when comparing to state-of-the-art methods. More remarkably, our MUSLE achieves to-date the best reported accuracy of 65.0% on Something-Something V2 validation set. * This work was performed at JD AI Research.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- Stand-Alone Inter-Frame Attention in Video ModelsFuchen Long, Zhaofan Qiu, Yingwei Pan, Ting Yao 等CVPR 2022 · 被引用 68 次
- Motion-Focused Contrastive Learning of Video Representations*Rui Li, Yiheng Zhang, Zhaofan Qiu, Ting Yao 等ICCV 2021 · 被引用 37 次
- MLP-3D: A MLP-like 3D Architecture with Grouped Time MixingZhaofan Qiu, Ting Yao, Chong-Wah Ngo, Tao MeiCVPR 2022 · 被引用 18 次
- Optimization Planning for 3D ConvNetsZhaofan Qiu, Ting Yao, Chong-Wah Ngo, Tao MeiICML 2021 · 被引用 9 次
- Reducing the Label Bias for Timestamp Supervised Temporal Action SegmentationKaiyuan Liu, Yunheng Li, Shenglan Liu, Chenwei Tan 等CVPR 2023
它引用的顶会 Paper6
- TSM: Temporal Shift Module for Efficient Video UnderstandingJi Lin, Chuang Gan, Song HanICCV 2019 · 被引用 2,049 次
- Reasoning About Human-Object Interactions Through Dual Attention NetworksTete Xiao, Quanfu Fan, Danny Gutfreund, Mathew Monfort 等ICCV 2019 · 被引用 36 次
- Region-Based Global Reasoning NetworksChuanming Wang, Huiyuan Fu, Charles X. Ling, Peilun Du 等AAAI 2020 · 被引用 5 次
- Adaptive Interaction Modeling via Graph Operations SearchHaoxin Li, Wei-Shi Zheng, Yu Tao, Haifeng Hu 等CVPR 2020
- Gate-Shift Networks for Video Action RecognitionSwathikiran Sudhakaran, Sergio Escalera, Oswald LanzCVPR 2020
相关 Paper
- Multi-Scale Spatial Temporal Graph Convolutional Network for Skeleton-Based Action RecognitionZhan Chen, Sicheng Li, Bing Yang, Qinghan Li 等AAAI 2021 · 被引用 341 次
- Multi-Group Multi-Attention: Towards Discriminative Spatiotemporal RepresentationZhensheng Shi, Liangjie Cao, Cheng Guan, Ju Liang 等ACM MM 2020 · 被引用 1 次
- SkeleTR: Towards Skeleton-based Action Recognition in the WildHaodong Duan, Mingze Xu, Bing Shuai, Davide Modolo 等ICCV 2023 · 被引用 38 次
- TubeR: Tubelet Transformer for Video Action DetectionJiaojiao Zhao, Yanyi Zhang, Xinyu Li, Hao Chen 等CVPR 2022 · 被引用 77 次
- Disentangling and Unifying Graph Convolutions for Skeleton-Based Action RecognitionZiyu Liu, Hongwen Zhang, Zhenghao Chen, Zhiyong Wang 等CVPR 2020
