Case Studies on the Motivation and Performance of Contributors Who Verify and Maintain In-Flux Tabular Datasets
Shaun Wallace, Alexandra Papoutsaki, Neilly H. Tan, Hua Guo, Jeff Huang
Abstract
The life cycle of a peer-produced dataset follows the phases of growth, maturity, and decline. Paying crowdworkers is a proven method to collect and organize information into structured tables. However, these tabular representations may contain inaccuracies due to errors or data changing over time. Thus, the maturation phase of a dataset can benefit from the additional human examination. One method to improve accuracy is to recruit additional paid crowdworkers to verify and correct errors. An alternative method relies on unpaid contributors, collectively editing the dataset during regular use. We describe two case studies to examine different strategies for human verification and maintenance of in-flux tabular datasets. The first case study examines traditional micro-task verification strategies with paid crowdworkers, while the second examines long-term maintenance strategies with unpaid contributions from non-crowdworkers. Two paid verification strategies that produced more accurate corrections at a lower cost per accurate correction were redundant data collection followed by final verification from a trusted crowdworker and allowing crowdworkers to review any data freely. In the unpaid maintenance strategies, contributors provided more accurate corrections when asked to review data matching their interests. This research identifies considerations and future approaches to collectively improving information accuracy and longevity of tabular information.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 78e804b3-3c91-404f-9f83-26829643e3b9Builds on3
- Keeping Community in the Loop: Understanding Wikipedia Stakeholder Values for Machine Learning-Based SystemsC. Estelle Smith, Bowen Yu, Anjali Srivastava, Aaron Halfaker et al.CHI 2020 · 77 citations
- Improving Worker Engagement Through Conversational Microtask CrowdsourcingSihang Qiu, Ujwal Gadiraju, Alessandro BozzonCHI 2020 · 66 citations
- Sketchy: Drawing Inspiration from the CrowdShaun Wallace, Brendan Le, Luis A. Leiva, Aman Haq et al.CSCW 2020 · 31 citations
Related papers
- Towards Fair and Equitable Incentives to Motivate Paid and Unpaid Crowd ContributionsShaun Wallace, Talie Massachi, Jiaqi Su, Dave Bryan Miller et al.CHI 2025 · 6 citations
- Productivity or Equity? Tradeoffs in Volunteer Microtasking in Humanitarian OpenStreetMapYaxuan Yin, Longjie Guo, Jacob Thebault-SpiekerCSCW 2024 · 6 citations
- Attract & Engage Visitors for Tabular Data Maintenance: a Longitudinal Naturalistic Field StudyShaun Wallace, Na Kyoung Lee, Zhengyi Peng, Talie Massachi et al.CSCW 2026
- Human-in-the-loop Regular Expression Extraction for Single Column Format InconsistencyShaochen Yu, Lei Han, Marta Indulska, Shazia Sadiq et al.WWW 2023 · 3 citations
- Crowdsourced Fact Validation for Knowledge BasesLibin Zheng, Peng Cheng, Lei Chen, Jianxing Yu et al.ICDE 2022 · 5 citations
