Regional Attention with Architecture-Rebuilt 3D Network for RGB-D Gesture Recognition
Benjia Zhou, Yunan Li, Jun Wan
摘要
Human gesture recognition has drawn much attention in the area of computer vision. However, the performance of gesture recognition is always influenced by some gesture-irrelevant factors like the background and the clothes of performers. Therefore, focusing on the regions of hand/arm is important to the gesture recognition. Meanwhile, a more adaptive architecture-searched network structure can also perform better than the block-fixed ones like ResNet since it increases the diversity of features in different stages of the network better. In this paper, we propose a regional attention with architecture-rebuilt 3D network (RAAR3DNet) for gesture recognition. We replace the fixed Inception modules with the automatically rebuilt structure through the network via Neural Architecture Search (NAS), owing to the different shape and representation ability of features in the early, middle, and late stage of the network. It enables the network to capture different levels of feature representations at different layers more adaptively. Meanwhile, we also design a stackable regional attention module called Dynamic-Static Attention (DSA), which derives a Gaussian guidance heatmap and dynamic motion map to highlight the hand/arm regions and the motion information in the spatial and temporal domains, respectively. Extensive experiments on two recent large-scale RGB-D gesture datasets validate the effectiveness of the proposed method and show it outperforms state-of-the-art methods. The codes of our method are available at: https://github.com/zhoubenjia/RAAR3DNet .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- Gloss-free Sign Language Translation: Improving from Visual-Language PretrainingBenjia Zhou, Zhigang Chen, Albert Clapés, Jun Wan 等ICCV 2023 · 被引用 123 次
- Nested Collaborative Learning for Long-Tailed Visual RecognitionJun Li, Zichang Tan, Jun Wan, Zhen Lei 等CVPR 2022 · 被引用 95 次
- Decoupling and Recoupling Spatiotemporal Representation for RGB-D-based Motion RecognitionBenjia Zhou, Pichao Wang, Jun Wan, Yanyan Liang 等CVPR 2022 · 被引用 43 次
- Multi-stage Factorized Spatio-Temporal Representation for RGB-D Action and Gesture RecognitionYujun Ma, Benjia Zhou, Ruili Wang, Pichao WangACM MM 2023 · 被引用 21 次
- Learning Robust Representations with Information Bottleneck and Memory Network for RGB-D-based Gesture RecognitionYunan Li, Huizhou Chen, Guanwen Feng, Qiguang MiaoICCV 2023 · 被引用 10 次
它引用的顶会 Paper2
相关 Paper
- Automatic Network Architecture Search for RGB-D Semantic SegmentationWenna Wang, Tao Zhuo, Xiuwei Zhang, Mingjun Sun 等ACM MM 2023 · 被引用 6 次
- Depth-Induced Multi-Scale Recurrent Attention Network for Saliency DetectionYongri Piao, Wei Ji, Jingjing Li, Miao Zhang 等ICCV 2019 · 被引用 450 次
- Dynamic Heterogeneous Graph Attention Neural Architecture SearchZeyang Zhang, Ziwei Zhang, Xin Wang, Yijian Qin 等AAAI 2023 · 被引用 44 次
- Dynamic Multiscale Graph Neural Networks for 3D Skeleton Based Human Motion PredictionMaosen Li, Siheng Chen, Yangheng Zhao, Ya Zhang 等CVPR 2020
- No Pain, Big Gain: Classify Dynamic Point Cloud Sequences with Static Models by Fitting Feature-level Space-time SurfacesJia-Xing Zhong, Kaichen Zhou, Qingyong Hu, Bing Wang 等CVPR 2022 · 被引用 22 次
