Exploring Collection of Sign Language Videos through Crowdsourcing
Danielle Bragg, Abraham Glasser, Fyodor Minakov, Naomi Caselli, William Thies
Abstract
Inadequate sign language data currently impedes advancement of sign language ML and AI. Training on existing datasets results in limited models due to small size, and lack of diverse signers in real-world settings. Complex labeling problems in particular often limit scale. In this work, we explore the potential for crowdsourcing to help overcome these barriers. To do this, we ran a user study with exploratory crowdsourcing tasks designed to support scalability: 1) to record videos of specific content -- thereby enabling automatic, scalable labeling -- and 2) to perform quality control checks for execution consistency -- further reducing post-processing requirements. We also provided workers with a searchable view of the crowdsourced dataset, to boost engagement and transparency and align with Deaf community values. Our user study included 29 participants using our exploratory tasks to record 1906 videos and perform 2331 quality control checks. Our results suggest that a crowd of signers may be able to generate high-quality recordings and perform reliable quality control, and that the signing community values visibility into the resulting dataset.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext f8201930-209a-4695-bc62-37b84e73dea9Cited by top-tier papers6
- Understanding the LLM-ification of CHI: Unpacking the Impact of LLMs at CHI through a Systematic Literature ReviewRock Yuren Pang, Hope Schroeder, Kynnedy Simone Smith, Solon Barocas et al.CHI 2025 · 51 citations
- Community-Driven Information Accessibility: Online Sign Language Content Creation within d/Deaf CommunitiesXinru Tang, Xiang Chang, Nuoran Chen, Yingjie (MaoMao) Ni et al.CHI 2023 · 21 citations
- Contributing to Accessibility Datasets: Reflections on Sharing Study Data by Blind PeopleRie Kamikubo, Kyungjun Lee, Hernisa KacorriCHI 2023 · 16 citations
- Disability-First AI Dataset Annotation: Co-designing Stuttered Speech Annotation Guidelines with People Who StutterXinru Tang, Jingjin Li, Shaomei WuCHI 2026 · 3 citations
- Reimagining Sign Language Technologies: Analyzing Translation Work of Chinese Deaf Online Content CreatorsXinru Tang, Anne Marie PiperCHI 2026 · 2 citations
Builds on2
- ASL Sea Battle: Gamifying Sign Language Data CollectionDanielle Bragg, Naomi Caselli, John W. Gallagher, Miriam Goldberg et al.CHI 2021 · 35 citations
- How2Sign: A Large-Scale Multimodal Dataset for Continuous American Sign LanguageAmanda Cardoso Duarte, Shruti Palaskar, Lucas Ventura, Deepti Ghadiyaram et al.CVPR 2021
Related papers
- Sign-Language Datasets at Scale: A Comprehensive Survey on Resources, Benchmarks, and Annotation StandardsYiming Ni, Zhi-Qi Cheng, Jiayu Li, Wei ChengACL 2026 · 1 citation
- Perceptions and Preferences: Deaf ASL-Signing Users' Insights on Video Elements, Styles and LayoutsKhulood Alkhudaidi, Tish Burke, Rachel Boll, Shruti Mahajan et al.CHI 2025 · 1 citation
- "I Hope This Is Helpful": Understanding Crowdworkers' Challenges and Motivations for an Image Description TaskRachel N. Simons, Danna Gurari, Kenneth R. FleischmannCSCW 2020 · 27 citations
- Towards AI-driven Sign Language Generation with Non-manual MarkersHan Zhang, Rotem Shalev-Arkushin, Vasileios Baltatzis, Connor Gillis et al.CHI 2025 · 12 citations
- YouTube-SL-25: A Large-Scale, Open-Domain Multilingual Sign Language Parallel CorpusGarrett Tanzer, Biao ZhangICLR 2025
