Exploring Collection of Sign Language Videos through Crowdsourcing
Danielle Bragg, Abraham Glasser, Fyodor Minakov, Naomi Caselli, William Thies
摘要
Inadequate sign language data currently impedes advancement of sign language ML and AI. Training on existing datasets results in limited models due to small size, and lack of diverse signers in real-world settings. Complex labeling problems in particular often limit scale. In this work, we explore the potential for crowdsourcing to help overcome these barriers. To do this, we ran a user study with exploratory crowdsourcing tasks designed to support scalability: 1) to record videos of specific content -- thereby enabling automatic, scalable labeling -- and 2) to perform quality control checks for execution consistency -- further reducing post-processing requirements. We also provided workers with a searchable view of the crowdsourced dataset, to boost engagement and transparency and align with Deaf community values. Our user study included 29 participants using our exploratory tasks to record 1906 videos and perform 2331 quality control checks. Our results suggest that a crowd of signers may be able to generate high-quality recordings and perform reliable quality control, and that the signing community values visibility into the resulting dataset.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- Understanding the LLM-ification of CHI: Unpacking the Impact of LLMs at CHI through a Systematic Literature ReviewRock Yuren Pang, Hope Schroeder, Kynnedy Simone Smith, Solon Barocas 等CHI 2025 · 被引用 51 次
- Community-Driven Information Accessibility: Online Sign Language Content Creation within d/Deaf CommunitiesXinru Tang, Xiang Chang, Nuoran Chen, Yingjie (MaoMao) Ni 等CHI 2023 · 被引用 21 次
- Contributing to Accessibility Datasets: Reflections on Sharing Study Data by Blind PeopleRie Kamikubo, Kyungjun Lee, Hernisa KacorriCHI 2023 · 被引用 16 次
- Disability-First AI Dataset Annotation: Co-designing Stuttered Speech Annotation Guidelines with People Who StutterXinru Tang, Jingjin Li, Shaomei WuCHI 2026 · 被引用 3 次
- Reimagining Sign Language Technologies: Analyzing Translation Work of Chinese Deaf Online Content CreatorsXinru Tang, Anne Marie PiperCHI 2026 · 被引用 2 次
它引用的顶会 Paper2
- ASL Sea Battle: Gamifying Sign Language Data CollectionDanielle Bragg, Naomi Caselli, John W. Gallagher, Miriam Goldberg 等CHI 2021 · 被引用 35 次
- How2Sign: A Large-Scale Multimodal Dataset for Continuous American Sign LanguageAmanda Cardoso Duarte, Shruti Palaskar, Lucas Ventura, Deepti Ghadiyaram 等CVPR 2021
相关 Paper
- Sign-Language Datasets at Scale: A Comprehensive Survey on Resources, Benchmarks, and Annotation StandardsYiming Ni, Zhi-Qi Cheng, Jiayu Li, Wei ChengACL 2026 · 被引用 1 次
- Perceptions and Preferences: Deaf ASL-Signing Users' Insights on Video Elements, Styles and LayoutsKhulood Alkhudaidi, Tish Burke, Rachel Boll, Shruti Mahajan 等CHI 2025 · 被引用 1 次
- "I Hope This Is Helpful": Understanding Crowdworkers' Challenges and Motivations for an Image Description TaskRachel N. Simons, Danna Gurari, Kenneth R. FleischmannCSCW 2020 · 被引用 27 次
- Towards AI-driven Sign Language Generation with Non-manual MarkersHan Zhang, Rotem Shalev-Arkushin, Vasileios Baltatzis, Connor Gillis 等CHI 2025 · 被引用 12 次
- YouTube-SL-25: A Large-Scale, Open-Domain Multilingual Sign Language Parallel CorpusGarrett Tanzer, Biao ZhangICLR 2025
