Studying and Mitigating Biases in Sign Language Understanding Models
Katherine Atwell, Danielle Bragg, Malihe Alikhani
Abstract
Ensuring that the benefits of sign language technologies are distributed equitably among all community members is crucial. Thus, it is important to address potential biases and inequities that may arise from the design or use of these resources. Crowd-sourced sign language datasets, such as the ASL Citizen dataset, are great resources for improving accessibility and preserving linguistic diversity, but they must be used thoughtfully to avoid reinforcing existing biases. In this work, we utilize the rich information about participant demographics and lexical features present in the ASL Citizen dataset to study and document the biases that may result from models trained on crowd-sourced sign datasets. Further, we apply several bias mitigation techniques during model training, and find that these techniques reduce performance disparities without decreasing accuracy. With the publication of this work, we release the demographic information about the participants in the ASL Citizen dataset to encourage future bias mitigation work in this space.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 68b201d2-44d2-47a0-b7f6-3d29c0dbfdbbCited by top-tier papers1
Ask how each one uses itBuilds on2
- OpenHands: Making Sign Language Recognition Accessible with Pose-based Pretrained Models across LanguagesPrem Selvaraj, Gokul N. C., Pratyush Kumar, Mitesh M. KhapraACL 2022 · 73 citations
- A Reduction to Binary Approach for Debiasing Multiclass DatasetsIbrahim M. Alabdulmohsin, Jessica Schrouff, Sanmi KoyejoNeurIPS 2022 · 11 citations
Related papers
- Exploring Collection of Sign Language Videos through CrowdsourcingDanielle Bragg, Abraham Glasser, Fyodor Minakov, Naomi Caselli et al.CSCW 2022 · 7 citations
- Impact of Annotator Demographics on Sentiment Dataset LabelingYi Ding, Jacob You, Tonja-Katrin Machulla, Jennifer Jacobs et al.CSCW 2022 · 20 citations
- Perturbation Augmentation for Fairer NLPRebecca Qian, Candace Ross, Jude Fernandes, Eric Michael Smith et al.EMNLP 2022 · 54 citations
- Sign-Language Datasets at Scale: A Comprehensive Survey on Resources, Benchmarks, and Annotation StandardsYiming Ni, Zhi-Qi Cheng, Jiayu Li, Wei ChengACL 2026 · 1 citation
- INCLUDE: A Large Scale Dataset for Indian Sign Language RecognitionAdvaith Sridhar, Rohith Gandhi Ganesan, Pratyush Kumar, Mitesh M. KhapraACM MM 2020 · 144 citations
