Wikibench: Community-Driven Data Curation for AI Evaluation on Wikipedia
Tzu-Sheng Kuo, Aaron Lee Halfaker, Zirui Cheng, Jiwoo Kim, Meng-Hsin Wu, Tongshuang Wu, Kenneth Holstein, Haiyi Zhu
摘要
AI tools are increasingly deployed in community contexts. However, datasets used to evaluate AI are typically created by developers and annotators outside a given community, which can yield misleading conclusions about AI performance. How might we empower communities to drive the intentional design and curation of evaluation datasets for AI that impacts them? We investigate this question on Wikipedia, an online community with multiple AI-based content moderation tools deployed. We introduce Wikibench, a system that enables communities to collaboratively curate AI evaluation datasets, while navigating ambiguities and differences in perspective through discussion. A field study on Wikipedia shows that datasets curated using Wikibench can effectively capture community consensus, disagreement, and uncertainty. Furthermore, study participants used Wikibench to shape the overall data curation process, including refining label definitions, determining data inclusion criteria, and authoring data statements. Based on our findings, we propose future directions for systems that support community-driven data curation.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper16
- Understanding the LLM-ification of CHI: Unpacking the Impact of LLMs at CHI through a Systematic Literature ReviewRock Yuren Pang, Hope Schroeder, Kynnedy Simone Smith, Solon Barocas 等CHI 2025 · 被引用 51 次
- Circuit Complexity Bounds for RoPE-based Transformer ArchitectureBo Chen, Xiaoyu Li, Yingyu Liang, Jiangxuan Long 等EMNLP 2025 · 被引用 33 次
- PolicyCraft: Supporting Collaborative and Participatory Policy Design through Case-Grounded DeliberationTzu-Sheng Kuo, Quan Ze Chen, Amy X. Zhang, Jane Hsieh 等CHI 2025 · 被引用 26 次
- Genius: A Generalizable and Purely Unsupervised Self-Training Framework For Advanced ReasoningFangzhi Xu, Hang Yan, Chang Ma, Haiteng Zhao 等ACL 2025 · 被引用 22 次
- WeAudit: Scaffolding User Auditors and AI Practitioners in Auditing Generative AIWesley Hanwen Deng, Claire Wang, Howard Ziyu Han, Jason I. Hong 等CSCW 2025 · 被引用 12 次
它引用的顶会 Paper25
- "Everyone wants to do the model work, not the data work": Data Cascades in High-Stakes AINithya Sambasivan, Shivani Kapania, Hannah Highfill, Diana Akrong 等CHI 2021 · 被引用 725 次
- Adversarial NLI: A New Benchmark for Natural Language UnderstandingYixin Nie, Adina Williams, Emily Dinan, Mohit Bansal 等ACL 2020 · 被引用 602 次
- Fine-tuning language models to find agreement among humans with diverse preferencesMichiel A. Bakker, Martin J. Chadwick, Hannah Sheahan, Michael Henry Tessler 等NeurIPS 2022 · 被引用 349 次
- Toward a Perspectivist Turn in Ground Truthing for Predictive ComputingFederico Cabitza, Andrea Campagner, Valerio BasileAAAI 2023 · 被引用 236 次
- Everyday Algorithm Auditing: Understanding the Power of Everyday Users in Surfacing Harmful Algorithmic BehaviorsHong Shen, Alicia DeVos, Motahhare Eslami, Kenneth HolsteinCSCW 2021 · 被引用 156 次
相关 Paper
- Engaging Communities Meaningfully in Defining Disability Representation for AI Image GenerationAnja Thieme, Rita Faia Marques, Martin Grayson, Sidhika Balachandar 等CHI 2026 · 被引用 1 次
- EditBench: Evaluating LLM Abilities to Perform Real-World Instructed Code EditsWayne Chi, Valerie Chen, Ryan Shar, Aditya Mittal 等ICLR 2026 · 被引用 7 次
- Computer-Aided Tagging on Wikimedia Commons: Designing for Human–AI Collaboration in Open Knowledge WorkYihan Yu, David W. McDonaldCSCW 2026
- The Human Labour of Data Work: Capturing Cultural Diversity through World Wide DishesSiobhan Mackenzie Hall, Samantha Dalal, Raesetje Sefala, Foutse Yuehgoh 等CSCW 2025 · 被引用 2 次
- UnsafeBench: Benchmarking Image Safety Classifiers on Real-World and AI-Generated ImagesYiting Qu, Xinyue Shen, Yixin Wu, Michael Backes 等CCS 2025 · 被引用 1 次
