Wikibench: Community-Driven Data Curation for AI Evaluation on Wikipedia
Tzu-Sheng Kuo, Aaron Lee Halfaker, Zirui Cheng, Jiwoo Kim, Meng-Hsin Wu, Tongshuang Wu, Kenneth Holstein, Haiyi Zhu
Abstract
AI tools are increasingly deployed in community contexts. However, datasets used to evaluate AI are typically created by developers and annotators outside a given community, which can yield misleading conclusions about AI performance. How might we empower communities to drive the intentional design and curation of evaluation datasets for AI that impacts them? We investigate this question on Wikipedia, an online community with multiple AI-based content moderation tools deployed. We introduce Wikibench, a system that enables communities to collaboratively curate AI evaluation datasets, while navigating ambiguities and differences in perspective through discussion. A field study on Wikipedia shows that datasets curated using Wikibench can effectively capture community consensus, disagreement, and uncertainty. Furthermore, study participants used Wikibench to shape the overall data curation process, including refining label definitions, determining data inclusion criteria, and authoring data statements. Based on our findings, we propose future directions for systems that support community-driven data curation.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b6e0a90a-8c88-4b2a-af3e-800f88da0e22Cited by top-tier papers16
- Understanding the LLM-ification of CHI: Unpacking the Impact of LLMs at CHI through a Systematic Literature ReviewRock Yuren Pang, Hope Schroeder, Kynnedy Simone Smith, Solon Barocas et al.CHI 2025 · 51 citations
- Circuit Complexity Bounds for RoPE-based Transformer ArchitectureBo Chen, Xiaoyu Li, Yingyu Liang, Jiangxuan Long et al.EMNLP 2025 · 33 citations
- PolicyCraft: Supporting Collaborative and Participatory Policy Design through Case-Grounded DeliberationTzu-Sheng Kuo, Quan Ze Chen, Amy X. Zhang, Jane Hsieh et al.CHI 2025 · 26 citations
- Genius: A Generalizable and Purely Unsupervised Self-Training Framework For Advanced ReasoningFangzhi Xu, Hang Yan, Chang Ma, Haiteng Zhao et al.ACL 2025 · 22 citations
- WeAudit: Scaffolding User Auditors and AI Practitioners in Auditing Generative AIWesley Hanwen Deng, Claire Wang, Howard Ziyu Han, Jason I. Hong et al.CSCW 2025 · 12 citations
Builds on25
- "Everyone wants to do the model work, not the data work": Data Cascades in High-Stakes AINithya Sambasivan, Shivani Kapania, Hannah Highfill, Diana Akrong et al.CHI 2021 · 725 citations
- Adversarial NLI: A New Benchmark for Natural Language UnderstandingYixin Nie, Adina Williams, Emily Dinan, Mohit Bansal et al.ACL 2020 · 602 citations
- Fine-tuning language models to find agreement among humans with diverse preferencesMichiel A. Bakker, Martin J. Chadwick, Hannah Sheahan, Michael Henry Tessler et al.NeurIPS 2022 · 349 citations
- Toward a Perspectivist Turn in Ground Truthing for Predictive ComputingFederico Cabitza, Andrea Campagner, Valerio BasileAAAI 2023 · 236 citations
- Everyday Algorithm Auditing: Understanding the Power of Everyday Users in Surfacing Harmful Algorithmic BehaviorsHong Shen, Alicia DeVos, Motahhare Eslami, Kenneth HolsteinCSCW 2021 · 156 citations
Related papers
- Engaging Communities Meaningfully in Defining Disability Representation for AI Image GenerationAnja Thieme, Rita Faia Marques, Martin Grayson, Sidhika Balachandar et al.CHI 2026 · 1 citation
- EditBench: Evaluating LLM Abilities to Perform Real-World Instructed Code EditsWayne Chi, Valerie Chen, Ryan Shar, Aditya Mittal et al.ICLR 2026 · 7 citations
- Computer-Aided Tagging on Wikimedia Commons: Designing for Human–AI Collaboration in Open Knowledge WorkYihan Yu, David W. McDonaldCSCW 2026
- The Human Labour of Data Work: Capturing Cultural Diversity through World Wide DishesSiobhan Mackenzie Hall, Samantha Dalal, Raesetje Sefala, Foutse Yuehgoh et al.CSCW 2025 · 2 citations
- UnsafeBench: Benchmarking Image Safety Classifiers on Real-World and AI-Generated ImagesYiting Qu, Xinyue Shen, Yixin Wu, Michael Backes et al.CCS 2025 · 1 citation
