Logos as a Well-Tempered Pre-train for Sign Language Recognition
Ilya Ovodov, Petr Surovtsev, Karina Kvanchiani, Alexander Kapitanov, Alexander Nagaev
摘要
This paper examines two aspects of the isolated sign language recognition (ISLR) task. First, although a certain number of datasets is available, the data for individual sign languages is limited. It poses the challenge of cross-language ISLR model training, including transfer learning. Second, similar signs can have different semantic meanings. It leads to ambiguity in dataset labeling and raises the question of the best policy for annotating such signs. To address these issues, this study presents Logos, a novel Russian Sign Language (RSL) dataset, the most extensive available ISLR dataset by the number of signers, one of the most extensive datasets in size and vocabulary, and the largest RSL dataset. It is shown that a model, pre-trained on the Logos dataset can be used as a universal encoder for other language SLR tasks, including few-shot learning. We explore cross-language transfer learning approaches and find that joint training using multiple classification heads benefits accuracy for the target low-resource datasets the most. The key feature of the Logos dataset is explicitly annotated visually similar sign groups. We show that explicitly labeling visually similar signs improves trained model quality as a visual encoder for downstream tasks. Based on the proposed contributions, we outperform current state-of-the-art results for the WLASL dataset and get competitive results for the AUTSL dataset, with a single stream model processing solely RGB video. The source code, dataset, and pre-trained models are publicly available.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper11
- CutMix: Regularization Strategy to Train Strong Classifiers With Localizable FeaturesSangdoo Yun, Dongyoon Han, Sanghyuk Chun, Seong Joon Oh 等ICCV 2019 · 被引用 5,843 次
- MViTv2: Improved Multiscale Vision Transformers for Classification and DetectionYanghao Li, Chao-Yuan Wu, Haoqi Fan, Karttikeya Mangalam 等CVPR 2022 · 被引用 699 次
- INCLUDE: A Large Scale Dataset for Indian Sign Language RecognitionAdvaith Sridhar, Rohith Gandhi Ganesan, Pratyush Kumar, Mitesh M. KhapraACM MM 2020 · 被引用 144 次
- A Simple Multi-Modality Transfer Learning Baseline for Sign Language TranslationYutong Chen, Fangyun Wei, Xiao Sun, Zhirong Wu 等CVPR 2022 · 被引用 137 次
- BEST: BERT Pre-training for Sign Language Recognition with Coupling TokenizationWeichao Zhao, Hezhen Hu, Wengang Zhou, Jiaxin Shi 等AAAI 2023 · 被引用 70 次
相关 Paper
- OpenHands: Making Sign Language Recognition Accessible with Pose-based Pretrained Models across LanguagesPrem Selvaraj, Gokul N. C., Pratyush Kumar, Mitesh M. KhapraACL 2022 · 被引用 73 次
- YouTube-SL-25: A Large-Scale, Open-Domain Multilingual Sign Language Parallel CorpusGarrett Tanzer, Biao ZhangICLR 2025
- SignCLIP: Connecting Text and Sign Language by Contrastive LearningZifan Jiang, Gerard Sant, Amit Moryossef, Mathias Müller 等EMNLP 2024 · 被引用 4 次
- Scaling Sign Language TranslationBiao Zhang, Garrett Tanzer, Orhan FiratNeurIPS 2024 · 被引用 21 次
- SignRep: Enhancing Self-Supervised Sign RepresentationsRyan Wong, Necati Cihan Camgöz, Richard BowdenICCV 2025 · 被引用 2 次
