"I Want to Publicize My Stutter": Community-led Collection and Curation of Chinese Stuttered Speech Data
Qisheng Li, Shaomei Wu
Abstract
This paper documents the process undertaken by StammerTalk , a grassroots community of Chinese-speaking people who stutter, to autonomously collect and curate stuttered speech data for more inclusive speech AI models. While people with disabilities are often excluded or treated merely as the subjects of AI data collection, our work introduces a new model for disability data collection in which the disability community exerts agency and control over their personal data and data-driven experiences. Our ethnographic data show that community-led data collection not only produces data needed to represent the community in AI systems, but also empowers the community and its members, by embracing - rather than concealing - stuttering and stutterer identity, and strengthening the social bonds of the community. Recognizing the lack of adequate socio-technical infrastructure for community-led, grassroots data collection, we discuss practical challenges, as well as the strategies and factors for communities to succeed in similar endeavors.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 750a0a35-ad2e-44e5-ba30-0a5b194a7d1bCited by top-tier papers3
- Disability-First AI Dataset Annotation: Co-designing Stuttered Speech Annotation Guidelines with People Who StutterXinru Tang, Jingjin Li, Shaomei WuCHI 2026 · 3 citations
- "It's trained by non-disabled people": Evaluating How Image Quality Affects Product Captioning with Vision-Language ModelsKapil Garg, Xinru Tang, Jimin Heo, Dwayne R. Morgan et al.CHI 2026 · 2 citations
- Engaging Communities Meaningfully in Defining Disability Representation for AI Image GenerationAnja Thieme, Rita Faia Marques, Martin Grayson, Sidhika Balachandar et al.CHI 2026 · 1 citation
Builds on3
- From User Perceptions to Technical Improvement: Enabling People Who Stutter to Better Use Speech RecognitionColin Lea, Zifang Huang, Jaya Narain, Lauren Tooley et al.CHI 2023 · 35 citations
- "It's not wrong, but I'm quite disappointed": Toward an Inclusive Algorithmic Experience for Content Creators with DisabilitiesDasom Choi, Uichin Lee, Hwajung HongCHI 2022 · 20 citations
- "The World is Designed for Fluent People": Benefits and Challenges of Videoconferencing Technologies for People Who StutterShaomei WuCHI 2023 · 15 citations
Related papers
- CoPracTter: Toward Integrating Personalized Practice Scenarios, Timely Feedback and Social Support into An Online Support Tool for Coping with Stuttering in ChinaLi Feng, Zeyu Xiong, Xinyi Li, Mingming FanCHI 2023 · 3 citations
- Finding My Voice over Zoom: An Autoethnography of Videoconferencing Experience for a Person Who StuttersShaomei Wu, Jingjin Li, Gilly LeshedCHI 2024 · 20 citations
- Situating Automatic Speech Recognition Development within Communities of Under-heard Language SpeakersThomas Reitmaier, Electra Wallington, Ondrej Klejch, Nina Markl et al.CHI 2023 · 9 citations
- ASTER: Automatic Speech Recognition System Accessibility Testing for StutterersYi Liu, Yuekang Li, Gelei Deng, Felix Juefei-Xu et al.ASE 2023 · 4 citations
- CS2W: A Chinese Spoken-to-Written Style Conversion Dataset with Multiple Conversion TypesZishan Guo, Linhao Yu, Minghui Xu, Renren Jin et al.EMNLP 2023
