BEST: BERT Pre-training for Sign Language Recognition with Coupling Tokenization
Weichao Zhao, Hezhen Hu, Wengang Zhou, Jiaxin Shi, Houqiang Li
Abstract
In this work, we are dedicated to leveraging the BERT pre-training success and modeling the domain-specific statistics to fertilize the sign language recognition (SLR) model. Considering the dominance of hand and body in sign language expression, we organize them as pose triplet units and feed them into the Transformer backbone in a frame-wise manner. Pre-training is performed via reconstructing the masked triplet unit from the corrupted input sequence, which learns the hierarchical correlation context cues among internal and external triplet units. Notably, different from the highly semantic word token in BERT, the pose unit is a low-level signal originally locating in continuous space, which prevents the direct adoption of the BERT cross entropy objective. To this end, we bridge this semantic gap via coupling tokenization of the triplet unit. It adaptively extracts the discrete pseudo label from the pose triplet unit, which represents the semantic gesture / body state. After pre-training, we fine-tune the pre-trained encoder on the downstream SLR task, jointly with the newly added task-specific layer. Extensive experiments are conducted to validate the effectiveness of our proposed method, achieving new state-of-the-art performance on all four benchmarks with a notable gain.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 8a05306c-9586-43d6-9b09-c7f9785c8fbaCited by top-tier papers9
- Geo-Sign: Hyperbolic Contrastive Regularisation for Geometrically Aware Sign Language TranslationEdward Fish, Richard BowdenNeurIPS 2025 · 15 citations
- Signs as Tokens: A Retrieval-Enhanced Multilingual Sign Language GeneratorRonglai Zuo, Rolandos Alexandros Potamias, Evangelos Ververas, Jiankang Deng et al.ICCV 2025 · 9 citations
- SignRep: Enhancing Self-Supervised Sign RepresentationsRyan Wong, Necati Cihan Camgöz, Richard BowdenICCV 2025 · 2 citations
- Logos as a Well-Tempered Pre-train for Sign Language RecognitionIlya Ovodov, Petr Surovtsev, Karina Kvanchiani, Alexander Kapitanov et al.EMNLP 2025 · 1 citation
- Natural Language-Assisted Sign Language RecognitionRonglai Zuo, Fangyun Wei, Brian MakCVPR 2023
Builds on16
- Zero-Shot Text-to-Image GenerationAditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray et al.ICML 2021 · 6,356 citations
- SlowFast Networks for Video RecognitionChristoph Feichtenhofer, Haoqi Fan, Jitendra Malik, Kaiming HeICCV 2019 · 4,104 citations
- BEiT: BERT Pre-Training of Image TransformersHangbo Bao, Li Dong, Songhao Piao, Furu WeiICLR 2022 · 3,632 citations
- TSM: Temporal Shift Module for Efficient Video UnderstandingJi Lin, Chuang Gan, Song HanICCV 2019 · 2,049 citations
- VideoBERT: A Joint Model for Video and Language Representation LearningChen Sun, Austin Myers, Carl Vondrick, Kevin Murphy et al.ICCV 2019 · 1,396 citations
Related papers
- SignBERT: Pre-Training of Hand-Model-Aware Representation for Sign Language RecognitionHezhen Hu, Weichao Zhao, Wengang Zhou, Yuechen Wang et al.ICCV 2021 · 125 citations
- SHuBERT: Self-Supervised Sign Language Representation Learning via Multi-Stream Cluster PredictionShester Gueuwou, Xiaodan Du, Greg Shakhnarovich, Karen Livescu et al.ACL 2025
- Uni-Sign: Toward Unified Sign Language Understanding at ScaleZecheng Li, Wengang Zhou, Weichao Zhao, Kepeng Wu et al.ICLR 2025
- Human Part-wise 3D Motion Context Learning for Sign Language RecognitionTaeryung Lee, Yeonguk Oh, Kyoung Mu LeeICCV 2023 · 14 citations
- Hand-Model-Aware Sign Language RecognitionHezhen Hu, Wengang Zhou, Houqiang LiAAAI 2021 · 79 citations
