LabelAId: Just-in-time AI Interventions for Improving Human Labeling Quality and Domain Knowledge in Crowdsourcing Systems
Chu Li, Zhihan Zhang, Michael Saugstad, Esteban Safranchik, Chaitanyashareef Kulkarni, Xiaoyu Huang, Shwetak N. Patel, Vikram Iyer, Tim Althoff, Jon E. Froehlich
Abstract
Crowdsourcing platforms have transformed distributed problem-solving, yet quality control remains a persistent challenge. Traditional quality control measures, such as prescreening workers and refining instructions, often focus solely on optimizing economic output. This paper explores just-in-time AI interventions to enhance both labeling quality and domain-specific knowledge among crowdworkers. We introduce LabelAId, an advanced inference model combining Programmatic Weak Supervision (PWS) with FT-Transformers to infer label correctness based on user behavior and domain knowledge. Our technical evaluation shows that our LabelAId pipeline consistently outperforms state-of-the-art ML baselines, improving mistake inference accuracy by 36.7% with 50 downstream samples. We then implemented LabelAId into Project Sidewalk, an open-source crowdsourcing platform for urban accessibility. A between-subjects study with 34 participants demonstrates that LabelAId significantly enhances label precision without compromising efficiency while also increasing labeler confidence. We discuss LabelAId’s success factors, limitations, and its generalizability to other crowdsourced science domains.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 6715e7c5-8d99-4145-8a76-47776617a157Cited by top-tier papers2
- Living Sustainability: In-Context Interactive Environmental Impact CommunicationZhihan Zhang, Puvarin Thavikulwat, Alexander Le Metzger, Yuxuan Mei et al.UbiComp 2025 · 5 citations
- CoKnowledge: Supporting Assimilation of Time-synced Collective Knowledge in Online Science VideosYuanhao Zhang, Yumeng Wang, Xiyuan Wang, Changyang He et al.CHI 2025 · 4 citations
Builds on11
- MLP-Mixer: An all-MLP Architecture for VisionIlya O. Tolstikhin, Neil Houlsby, Alexander Kolesnikov, Lucas Beyer et al.NeurIPS 2021 · 3,862 citations
- Revisiting Deep Learning Models for Tabular DataYury Gorishniy, Ivan Rubachev, Valentin Khrulkov, Artem BabenkoNeurIPS 2021 · 1,847 citations
- To Trust or to Think: Cognitive Forcing Functions Can Reduce Overreliance on AI in AI-assisted Decision-makingZana Buçinca, Maja Barbara Malaya, Krzysztof Z. GajosCSCW 2021 · 962 citations
- Does the Whole Exceed its Parts? The Effect of AI Explanations on Complementary Team PerformanceGagan Bansal, Tongshuang Wu, Joyce Zhou, Raymond Fok et al.CHI 2021 · 713 citations
- On the Use of Multi-sensory Cues in Symmetric and Asymmetric Shared Collaborative Virtual SpacesSungchul Jung, Nawam Karki, Max W. J. Slutter, Robert W. LindemanCSCW 2021 · 33 citations
Related papers
- "I never realized sidewalks were a big deal": A Case Study of a Community-Driven Sidewalk Accessibility Assessment using Project SidewalkChu Li, Katrina Oi Yau Ma, Michael Saugstad, Kie Fujii et al.CHI 2024 · 12 citations
- Refining Labeling Functions with Limited Labeled DataChenjie Li, Amir Gilad, Boris Glavic, Zhengjie Miao et al.KDD 2025 · 1 citation
- Learning Hyper Label Model for Programmatic Weak SupervisionRenzhi Wu, Shen-En Chen, Jieyu Zhang, Xu ChuICLR 2023 · 2 citations
- Weak Supervision Performance Evaluation via Partial IdentificationFelipe Maia Polo, Subha Maity, Mikhail Yurochkin, Moulinath Banerjee et al.NeurIPS 2024 · 6 citations
- Active Label Correction for Semantic Segmentation with Foundation ModelsHoyoung Kim, Sehyun Hwang, Suha Kwak, Jungseul OkICML 2024 · 5 citations
