CalCo: A Hierarchical Bayesian Framework for Scalable Human-LLM Hybrid Labeling
Viet-An Nguyen, Xu Chen, Udi Weinsberg
Abstract
With the emergence of Large Language Models (LLMs), data annotation workflows are increasingly shifting toward hybrid systems where LLM-generated labels complement or replace human effort. However, effectively incorporating these non-deterministic signals into existing human labeling processes remains a challenge, particularly in complex industrial settings involving multiple correlated tasks and diverse model architectures. In this work, we introduce CalCo, a comprehensive system deployed at scale, designed to integrate LLM and human labels in a principled manner. At the core of CalCo is a novel hierarchical Bayesian model that explicitly captures the complex dependencies inherent in the multi-task, multi-LLM annotation process. Unlike traditional approaches that treat annotation inputs in isolation, CalCo jointly models the inter-dependencies across related tasks and the correlations among diverse LLMs. This holistic approach allows the system to share statistical strength across questions and capture shared error profiles among models. We demonstrate the efficacy of our approach through extensive empirical results on both simulated and real-world datasets, showing that CalCo significantly outperforms baseline methods in a wide range of applications.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 1e7715f7-589d-43fa-8e00-9ba3eea9702aRelated papers
- Human-LLM Collaborative Annotation Through Effective Verification of LLM LabelsXinru Wang, Hannah Kim, Sajjadur Rahman, Kushan Mitra et al.CHI 2024 · 127 citations
- Label Annotation for Tabular Anomaly Detection with Large Language ModelsHaihong Zhao, Aochuan Chen, Miao Peng, Xiaolong Fan et al.KDD 2026
- Multicalibration for Confidence Scoring in LLMsGianluca Detommaso, Martin Bertran Lopez, Riccardo Fogliato, Aaron RothICML 2024 · 39 citations
- Knowledge Exchange with Confidence: Cost-Effective LLM Integration for Reliable and Efficient Visual Question AnsweringMahsa Mozaffari, Hitesh Sapkota, Xumin Liu, Qi YuICLR 2026
- Next Generation Active Learning: Mixture of LLMs in the LoopYuanyuan Qi, Xiaohao Yang, Jueqing Lu, Guoxiang Guo et al.AAAI 2026
