Human Part-wise 3D Motion Context Learning for Sign Language Recognition
Taeryung Lee, Yeonguk Oh, Kyoung Mu Lee
Abstract
In this paper, we propose P3D, the human part-wise motion context learning framework for sign language recognition. Our main contributions lie in two dimensions: learning the part-wise motion context and employing the pose ensemble to utilize 2D and 3D pose jointly. First, our empirical observation implies that part-wise context encoding benefits the performance of sign language recognition. While previous methods of sign language recognition learned motion context from the sequence of the entire pose, we argue that such methods cannot exploit part-specific motion context. In order to utilize part-wise motion context, we propose the alternating combination of a part-wise encoding Transformer (PET) and a whole-body encoding Transformer (WET). PET encodes the motion contexts from a part sequence, while WET merges them into a unified context. By learning part-wise motion context, our P3D achieves superior performance on WLASL compared to previous state-of-the-art methods. Second, our framework is the first to ensemble 2D and 3D poses for sign language recognition. Since the 3D pose holds rich motion context and depth information to distinguish the words, our P3D outperformed the previous state-of-the-art methods employing a pose ensemble.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext a8bea89e-7c28-47ae-ab77-b8f8f50fe449Cited by top-tier papers5
- Towards Online Continuous Sign Language Recognition and TranslationRonglai Zuo, Fangyun Wei, Brian MakEMNLP 2024 · 14 citations
- SEDS: Semantically Enhanced Dual-Stream Encoder for Sign Language RetrievalLongtao Jiang, Min Wang, Zecheng Li, Yao Fang et al.ACM MM 2024 · 2 citations
- Cross-View Isolated Sign Language Recognition via View Synthesis and Feature DisentanglementXin Shen, Xinyu Wang, Lei Shen, Kaihao Zhang et al.ICCV 2025 · 1 citation
- T2S-GPT: Dynamic Vector Quantization for Autoregressive Sign Language Production from TextAoxiong Yin, Haoyuan Li, Kai Shen, Siliang Tang et al.ACL 2024
- VSNet: Focusing on the Linguistic Characteristics of Sign LanguageYuhao Li, Xinyue Chen, Hongkai Li, Xiaorong Pu et al.CVPR 2025
Builds on13
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- ViViT: A Video Vision TransformerAnurag Arnab, Mostafa Dehghani, Georg Heigold, Chen Sun et al.ICCV 2021 · 2,947 citations
- Revisiting Skeleton-based Action RecognitionHaodong Duan, Yue Zhao, Kai Chen, Dahua Lin et al.CVPR 2022 · 752 citations
- 3D Human Pose Estimation with Spatial and Temporal TransformersCe Zheng, Sijie Zhu, Matías Mendieta, Taojiannan Yang et al.ICCV 2021 · 648 citations
- MHFormer: Multi-Hypothesis Transformer for 3D Human Pose EstimationWenhao Li, Hong Liu, Hao Tang, Pichao Wang et al.CVPR 2022 · 403 citations
Related papers
- BEST: BERT Pre-training for Sign Language Recognition with Coupling TokenizationWeichao Zhao, Hezhen Hu, Wengang Zhou, Jiaxin Shi et al.AAAI 2023 · 70 citations
- Hand-Model-Aware Sign Language RecognitionHezhen Hu, Wengang Zhou, Houqiang LiAAAI 2021 · 79 citations
- SignBERT: Pre-Training of Hand-Model-Aware Representation for Sign Language RecognitionHezhen Hu, Weichao Zhao, Wengang Zhou, Yuechen Wang et al.ICCV 2021 · 125 citations
- SignRep: Enhancing Self-Supervised Sign RepresentationsRyan Wong, Necati Cihan Camgöz, Richard BowdenICCV 2025 · 2 citations
- Skeleton-Aware Neural Sign Language TranslationShiwei Gan, Yafeng Yin, Zhiwei Jiang, Lei Xie et al.ACM MM 2021 · 28 citations
