The Human Labour of Data Work: Capturing Cultural Diversity through World Wide Dishes
Siobhan Mackenzie Hall, Samantha Dalal, Raesetje Sefala, Foutse Yuehgoh, Aisha Alaagib, Imane Hamzaoui, Shu Ishida, Jabez Magomere, Lauren Crais, Aya Salama, Tejumade Afonja
Abstract
This paper provides guidance for building and maintaining infrastructure for participatory AI efforts by sharing reflections on building World Wide Dishes (WWD), a bottom-up, community-led image and text dataset of culinary dishes and associated cultural customs. We present WWD as an example of participatory dataset creation, where community members both guide the design of the research process and contribute to the crowdsourced dataset. This approach incorporates localised expertise and knowledge to address the limitations of web-scraped Internet datasets acknowledged in the Participatory AI discourse. We show that our approach can result in curated, high-quality data that supports decentralised contributions from communities that do not typically contribute to datasets due to a variety of systemic factors. Our project demonstrates the importance of participatory mediators in supporting community engagement by identifying the kinds of labour they performed to make WWD possible. We surface three dimensions of labour performed by participatory mediators that are crucial for participatory dataset construction: building trust with community members, making participation accessible, and contextualising community values to support meaningful data collection. Drawing on our findings, we put forth five lessons for building infrastructure to support future participatory AI efforts.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers2
- Building Benchmarks from the Ground Up: Community-Centered Evaluation of LLMs in Healthcare Chatbot SettingsHamna, Gayatri Bhat, Sourabrata Mukherjee, Faisal M. Lalani et al.CHI 2026 · 1 citation
- Engaging Communities Meaningfully in Defining Disability Representation for AI Image GenerationAnja Thieme, Rita Faia Marques, Martin Grayson, Sidhika Balachandar et al.CHI 2026 · 1 citation
Builds on20
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- "Everyone wants to do the model work, not the data work": Data Cascades in High-Stakes AINithya Sambasivan, Shivani Kapania, Hannah Highfill, Diana Akrong et al.CHI 2021 · 725 citations
- Data Filtering NetworksAlex Fang, Albin Madappally Jose, Amit Jain, Ludwig Schmidt et al.ICLR 2024 · 251 citations
- Do Datasets Have Politics? Disciplinary Values in Computer Vision Dataset DevelopmentMorgan Klaus Scheuerman, Alex Hanna, Emily DentonCSCW 2021 · 169 citations
- I Can't Breathe: Reflections from Black Women in CSCW and HCISheena Lewis Erete, Yolanda A. Rankin, Jakita Owensby ThomasCSCW 2020 · 122 citations
Related papers
- Wikibench: Community-Driven Data Curation for AI Evaluation on WikipediaTzu-Sheng Kuo, Aaron Lee Halfaker, Zirui Cheng, Jiwoo Kim et al.CHI 2024 · 19 citations
- What are Foundation Models Cooking in the Post-Soviet World?Anton Lavrouk, Tarek Naous, Alan Ritter, Wei XuEMNLP 2025
- Water On My Block: Reflections on Building A Participatory Artificial Intelligence System For Precision Weather With Scientists and An Urban CommunityKelly B. Wagman, Kanchan Uday Naik, Madison Vanderbilt, Naeun Ko et al.CHI 2026
- ViGiL3D: A Linguistically Diverse Dataset for 3D Visual GroundingAustin T. Wang, ZeMing Gong, Angel X. ChangACL 2025 · 6 citations
- Aya Dataset: An Open-Access Collection for Multilingual Instruction TuningShivalika Singh, Freddie Vargus, Daniel D'souza, Börje Karlsson et al.ACL 2024
