INCLUDE: A Large Scale Dataset for Indian Sign Language Recognition
Advaith Sridhar, Rohith Gandhi Ganesan, Pratyush Kumar, Mitesh M. Khapra
Abstract
Indian Sign Language (ISL) is a complete language with its own grammar, syntax, vocabulary and several unique linguistic attributes. It is used by over 5 million deaf people in India. Currently, there is no publicly available dataset on ISL to evaluate Sign Language Recognition (SLR) approaches. In this work, we present the Indian Lexicon Sign Language Dataset - INCLUDE - an ISL dataset that contains 0.27 million frames across 4,287 videos over 263 word signs from 15 different word categories. INCLUDE is recorded with the help of experienced signers to provide close resemblance to natural conditions. A subset of 50 word signs is chosen across word categories to define INCLUDE-50 for rapid evaluation of SLR meth- ods with hyperparameter tuning. As the first large scale study of SLR on ISL, we evaluate several deep neural networks combining different methods for augmentation, feature extraction, encoding and decoding. The best performing model achieves an accuracy of 94.5% on the INCLUDE-50 dataset and 85.6% on the INCLUDE dataset. This model uses a pre-trained feature extractor and encoder and only trains a decoder. We further explore generalisation by fine-tuning the decoder for an American Sign Language dataset. On the ASLLVD with 48 classes, our model has an accuracy of 92.1%; improving on existing results and providing an efficient method to support SLR for multiple languages.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 1ecacf60-57ca-45fb-b632-64cb5e7a90e8Cited by top-tier papers6
- SCOPE: Sign Language Contextual Processing with Embedding from LLMsYuqi Liu, Wenqian Zhang, Sihan Ren, Chengyu Huang et al.AAAI 2025 · 7 citations
- Logos as a Well-Tempered Pre-train for Sign Language RecognitionIlya Ovodov, Petr Surovtsev, Karina Kvanchiani, Alexander Kapitanov et al.EMNLP 2025 · 1 citation
- Sign-Language Datasets at Scale: A Comprehensive Survey on Resources, Benchmarks, and Annotation StandardsYiming Ni, Zhi-Qi Cheng, Jiayu Li, Wei ChengACL 2026 · 1 citation
- Improving Sign Language Translation With Monolingual Data by Sign Back-TranslationHao Zhou, Wengang Zhou, Weizhen Qi, Junfu Pu et al.CVPR 2021
- BhashaSutra: A Task-Centric Unified Survey of Indian NLP Datasets, Corpora, and ResourcesRaghvendra Kumar, Devankar Raj, Sriparna SahaACL 2026
Related papers
- YouTube-SL-25: A Large-Scale, Open-Domain Multilingual Sign Language Parallel CorpusGarrett Tanzer, Biao ZhangICLR 2025
- Reconstructing Signing Avatars from Video Using Linguistic PriorsMaria-Paola Forte, Peter Kulits, Chun-Hao Huang, Vasileios Choutas et al.CVPR 2023
- SignQuery: A Natural User Interface and Search Engine for Sign Languages with Wearable SensorsHao Zhou, Taiting Lu, Kristina Mckinnie, Joseph Palagano et al.MobiCom 2023 · 13 citations
- Open-Domain Sign Language Translation Learned from Online VideoBowen Shi, Diane Brentari, Gregory Shakhnarovich, Karen LivescuEMNLP 2022 · 39 citations
- OpenHands: Making Sign Language Recognition Accessible with Pose-based Pretrained Models across LanguagesPrem Selvaraj, Gokul N. C., Pratyush Kumar, Mitesh M. KhapraACL 2022 · 73 citations
