Disability-First AI Dataset Annotation: Co-designing Stuttered Speech Annotation Guidelines with People Who Stutter
Xinru Tang, Jingjin Li, Shaomei Wu
Abstract
Despite efforts to increase the representation of disabled people in AI datasets, accessibility datasets are often annotated by crowdworkers without disability-specific expertise, leading to inconsistent or inaccurate labels. This paper examines these annotation challenges through a case study of annotating speech data from people who stutter (PWS). Given the variability of stuttering and differing views on how it manifests, annotating and transcribing stuttered speech remains difficult, even for trained professionals. Through interviews and co-design workshops with PWS and domain experts, we identify challenges in stuttered speech annotation and develop practices that integrate the lived experiences of PWS into the annotation process. Our findings highlight the value of embodied knowledge in improving dataset quality, while revealing tensions between the complexity of disability experiences and the rigidity of static labels. We conclude with implications for disability-first and multiplicity-aware approaches to data interpretation across the AI pipeline.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext c51cf20b-cd16-464b-9d51-67b142cd84c7Cited by top-tier papers2
- Access in the Shadow of Ableism: An Autoethnography of a Blind Student's Higher Education Experience in ChinaXinru Tang, Weijun ZhangCHI 2026 · 3 citations
- Reimagining Sign Language Technologies: Analyzing Translation Work of Chinese Deaf Online Content CreatorsXinru Tang, Anne Marie PiperCHI 2026 · 2 citations
Builds on31
- Robust Speech Recognition via Large-Scale Weak SupervisionAlec Radford, Jong Wook Kim, Tao Xu, Greg Brockman et al.ICML 2023 · 6,966 citations
- How We've Taught Algorithms to See Identity: Constructing Race and Gender in Image Databases for Facial AnalysisMorgan Klaus Scheuerman, Kandrea Wade, Caitlin Lustig, Jed R. BrubakerCSCW 2020 · 198 citations
- Between Subjectivity and Imposition: Power Dynamics in Data Annotation for Computer VisionMilagros Miceli, Martin Schuessler, Tianling YangCSCW 2020 · 148 citations
- "It's Complicated": Negotiating Accessibility and (Mis)Representation in Image Descriptions of Race, Gender, and DisabilityCynthia L. Bennett, Cole Gleason, Morgan Klaus Scheuerman, Jeffrey P. Bigham et al.CHI 2021 · 125 citations
- Designing Ground Truth and the Social Life of LabelsMichael J. Muller, Christine T. Wolf, Josh Andres, Michael Desmond et al.CHI 2021 · 80 citations
Related papers
- "I Want to Publicize My Stutter": Community-led Collection and Curation of Chinese Stuttered Speech DataQisheng Li, Shaomei WuCSCW 2024 · 3 citations
- Finding My Voice over Zoom: An Autoethnography of Videoconferencing Experience for a Person Who StuttersShaomei Wu, Jingjin Li, Gilly LeshedCHI 2024 · 20 citations
- ASTER: Automatic Speech Recognition System Accessibility Testing for StutterersYi Liu, Yuekang Li, Gelei Deng, Felix Juefei-Xu et al.ASE 2023 · 4 citations
- "Computer Says No": Disabled Welfare Experiences and Envisioned Futures Under AI GovernanceHumphrey Curtis, Adam D. G. Jenkins, Alistair Gentry, Sioban Zacharek et al.CHI 2026 · 2 citations
- From User Perceptions to Technical Improvement: Enabling People Who Stutter to Better Use Speech RecognitionColin Lea, Zifang Huang, Jaya Narain, Lauren Tooley et al.CHI 2023 · 35 citations
