Siformer: Feature-isolated Transformer for Efficient Skeleton-based Sign Language Recognition
Muxin Pu, Mei Kuan Lim, Chun Yong Chong
Abstract
Sign language recognition (SLR) refers to interpreting sign language glosses from given videos automatically. This research area presents a complex challenge in computer vision because of the rapid and intricate movements inherent in sign languages, which encompass hand gestures, body postures, and even facial expressions. Recently, skeleton-based action recognition has attracted increasing attention due to its ability to handle variations in subjects and backgrounds independently. However, current skeleton-based SLR methods exhibit three limitations: 1) they often neglect the importance of realistic hand poses, where most studies train SLR models on non-realistic skeletal representations; 2) they tend to assume complete data availability in both training or inference phases, and capture intricate relationships among different body parts collectively; 3) these methods treat all sign glosses uniformly, failing to account for differences in complexity levels regarding skeletal representations. To enhance the realism of hand skeletal representations, we present a kinematic hand pose rectification method for enforcing constraints. Mitigating the impact of missing data, we propose a feature-isolated mechanism to focus on capturing local spatial-temporal context. This method captures the context concurrently and independently from individual features, thus enhancing the robustness of the SLR model. Additionally, to adapt to varying complexity levels of sign glosses, we develop an inputadaptive inference approach to optimise computational efficiency and accuracy. Experimental results demonstrate the effectiveness of our approach, as evidenced by achieving a new state-of-the-art (SOTA) performance on WLASL100 and LSA64. For WLASL100, we achieve a top-1 accuracy of 86.50%, marking a relative improvement of 2.39% over the previous SOTA. For LSA64, we achieve a top-1 accuracy of 99.84%. The artefacts and code related to this study are made publicly online. 1
• Computing methodologies → Supervised learning by classification; Interest point and salient region detections; Feature selection; Neural networks; • Human-centered computing → Human computer interaction (HCI).
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 89cf279c-b255-478b-95a9-90abf32eafd4Cited by top-tier papers1
Ask how each one uses itBuilds on8
- Informer: Beyond Efficient Transformer for Long Sequence Time-Series ForecastingHaoyi Zhou, Shanghang Zhang, Jieqi Peng, Shuai Zhang et al.AAAI 2021 · 7,289 citations
- Reducing Transformer Depth on Demand with Structured DropoutAngela Fan, Edouard Grave, Armand JoulinICLR 2020 · 695 citations
- BERT Loses Patience: Fast and Robust Inference with Early ExitWangchunshu Zhou, Canwen Xu, Tao Ge, Julian J. McAuley et al.NeurIPS 2020 · 473 citations
- BERT-of-Theseus: Compressing BERT by Progressive Module ReplacingCanwen Xu, Wangchunshu Zhou, Tao Ge, Furu Wei et al.EMNLP 2020 · 168 citations
- Triple Wins: Boosting Accuracy, Robustness and Efficiency Together by Enabling Input-Adaptive InferenceTing-Kuei Hu, Tianlong Chen, Haotao Wang, Zhangyang WangICLR 2020 · 89 citations
Related papers
- VSNet: Focusing on the Linguistic Characteristics of Sign LanguageYuhao Li, Xinyue Chen, Hongkai Li, Xiaorong Pu et al.CVPR 2025
- Hand-Model-Aware Sign Language RecognitionHezhen Hu, Wengang Zhou, Houqiang LiAAAI 2021 · 79 citations
- SignBERT: Pre-Training of Hand-Model-Aware Representation for Sign Language RecognitionHezhen Hu, Weichao Zhao, Wengang Zhou, Yuechen Wang et al.ICCV 2021 · 125 citations
- Natural Language-Assisted Sign Language RecognitionRonglai Zuo, Fangyun Wei, Brian MakCVPR 2023
- SignRep: Enhancing Self-Supervised Sign RepresentationsRyan Wong, Necati Cihan Camgöz, Richard BowdenICCV 2025 · 2 citations
