Human Part-wise 3D Motion Context Learning for Sign Language Recognition
Taeryung Lee, Yeonguk Oh, Kyoung Mu Lee
摘要
In this paper, we propose P3D, the human part-wise motion context learning framework for sign language recognition. Our main contributions lie in two dimensions: learning the part-wise motion context and employing the pose ensemble to utilize 2D and 3D pose jointly. First, our empirical observation implies that part-wise context encoding benefits the performance of sign language recognition. While previous methods of sign language recognition learned motion context from the sequence of the entire pose, we argue that such methods cannot exploit part-specific motion context. In order to utilize part-wise motion context, we propose the alternating combination of a part-wise encoding Transformer (PET) and a whole-body encoding Transformer (WET). PET encodes the motion contexts from a part sequence, while WET merges them into a unified context. By learning part-wise motion context, our P3D achieves superior performance on WLASL compared to previous state-of-the-art methods. Second, our framework is the first to ensemble 2D and 3D poses for sign language recognition. Since the 3D pose holds rich motion context and depth information to distinguish the words, our P3D outperformed the previous state-of-the-art methods employing a pose ensemble.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- Towards Online Continuous Sign Language Recognition and TranslationRonglai Zuo, Fangyun Wei, Brian MakEMNLP 2024 · 被引用 14 次
- SEDS: Semantically Enhanced Dual-Stream Encoder for Sign Language RetrievalLongtao Jiang, Min Wang, Zecheng Li, Yao Fang 等ACM MM 2024 · 被引用 2 次
- Cross-View Isolated Sign Language Recognition via View Synthesis and Feature DisentanglementXin Shen, Xinyu Wang, Lei Shen, Kaihao Zhang 等ICCV 2025 · 被引用 1 次
- T2S-GPT: Dynamic Vector Quantization for Autoregressive Sign Language Production from TextAoxiong Yin, Haoyuan Li, Kai Shen, Siliang Tang 等ACL 2024
- VSNet: Focusing on the Linguistic Characteristics of Sign LanguageYuhao Li, Xinyue Chen, Hongkai Li, Xiaorong Pu 等CVPR 2025
它引用的顶会 Paper13
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- ViViT: A Video Vision TransformerAnurag Arnab, Mostafa Dehghani, Georg Heigold, Chen Sun 等ICCV 2021 · 被引用 2,947 次
- Revisiting Skeleton-based Action RecognitionHaodong Duan, Yue Zhao, Kai Chen, Dahua Lin 等CVPR 2022 · 被引用 752 次
- 3D Human Pose Estimation with Spatial and Temporal TransformersCe Zheng, Sijie Zhu, Matías Mendieta, Taojiannan Yang 等ICCV 2021 · 被引用 648 次
- MHFormer: Multi-Hypothesis Transformer for 3D Human Pose EstimationWenhao Li, Hong Liu, Hao Tang, Pichao Wang 等CVPR 2022 · 被引用 403 次
相关 Paper
- BEST: BERT Pre-training for Sign Language Recognition with Coupling TokenizationWeichao Zhao, Hezhen Hu, Wengang Zhou, Jiaxin Shi 等AAAI 2023 · 被引用 70 次
- Hand-Model-Aware Sign Language RecognitionHezhen Hu, Wengang Zhou, Houqiang LiAAAI 2021 · 被引用 79 次
- SignBERT: Pre-Training of Hand-Model-Aware Representation for Sign Language RecognitionHezhen Hu, Weichao Zhao, Wengang Zhou, Yuechen Wang 等ICCV 2021 · 被引用 125 次
- SignRep: Enhancing Self-Supervised Sign RepresentationsRyan Wong, Necati Cihan Camgöz, Richard BowdenICCV 2025 · 被引用 2 次
- Skeleton-Aware Neural Sign Language TranslationShiwei Gan, Yafeng Yin, Zhiwei Jiang, Lei Xie 等ACM MM 2021 · 被引用 28 次
