LANCET: Labeling Complex Data at Scale
Huayi Zhang, Lei Cao, Samuel Madden, Elke A. Rundensteiner
Abstract
Cutting-edge machine learning techniques often require millions of labeled data objects to train a robust model. Because relying on humans to supply such a huge number of labels is rarely practical, automated methods for label generation are needed. Unfortunately, critical challenges in auto-labeling remain unsolved, including the following research questions: (1) which objects to ask humans to label, (2) how to automatically propagate labels to other objects, and (3) when to stop labeling. These three questions are not only each challenging in their own right, but they also correspond to tightly interdependent problems. Yet existing techniques provide at best isolated solutions to a subset of these challenges. In this work, we propose the first approach, called LANCET, that successfully addresses all three challenges in an integrated framework. LANCET is based on a theoretical foundation characterizing the properties that the labeled dataset must satisfy to train an effective prediction model, namely the Covariate-shift and the Continuity conditions. First, guided by the Covariate-shift condition, LANCET maps raw input data into a semantic feature space, where an unlabeled object is expected to share the same label with its near-by labeled neighbor. Next, guided by the Continuity condition, LANCET selects objects for labeling, aiming to ensure that unlabeled objects always have some sufficiently close labeled neighbors. These two strategies jointly maximize the accuracy of the automatically produced labels and the prediction accuracy of the machine learning models trained on these labels. Lastly, LANCET uses a distribution matching network to verify whether both the Covariate-shift and Continuity conditions hold, in which case it would be safe to terminate the labeling process. Our experiments on diverse public data sets demonstrate that LANCET consistently outperforms the state-of-the-art methods from Snuba to GOGGLES and other baselines by a large margin - up to 30 percentage points increase in accuracy.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 14576e57-0f62-47f9-9300-98fa3218e997Cited by top-tier papers4
- MetaStore: Analyzing Deep Learning Meta-Data at ScaleHuayi Zhang, Binwei Yan, Lei Cao, Samuel Madden et al.VLDB 2024 · 10 citations
- VOCALExplore: Pay-as-You-Go Video Data Exploration and Model BuildingMaureen Daum, Enhao Zhang, Dong He, Stephen Mussmann et al.VLDB 2023 · 7 citations
- CLaDMoP: Learning Transferrable Models from Successful Clinical Trials via LLMsYiqing Zhang, Xiaozhong Liu, Fabricio MuraiKDD 2025 · 1 citation
- CLID-MU: Cross-Layer Information Divergence Based Meta Update Strategy for Learning with Noisy LabelsRuofan Hu, Dongyu Zhang, Huayi Zhang, Elke A. RundensteinerKDD 2025
Builds on4
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- FixMatch: Simplifying Semi-Supervised Learning with Consistency and ConfidenceKihyuk Sohn, David Berthelot, Nicholas Carlini, Zizhao Zhang et al.NeurIPS 2020 · 5,129 citations
- Variational Adversarial Active LearningSamarth Sinha, Sayna Ebrahimi, Trevor DarrellICCV 2019 · 662 citations
- GOGGLES: Automatic Image Labeling with Affinity CodingNilaksh Das, Sanya Chaba, Renzhi Wu, Sakshi Gandhi et al.SIGMOD 2020 · 22 citations
Related papers
- Pearls from Pebbles: Improved Confidence Functions for Auto-labelingHarit Vishwakarma, Yi Chen, Sui Jiet Tay, Satya Sai Srinath Namburi et al.NeurIPS 2024 · 7 citations
- Feed Two Birds with One Scone: Exploiting Wild Data for Both Out-of-Distribution Generalization and DetectionHaoyue Bai, Gregory Canal, Xuefeng Du, Jeongyeol Kwon et al.ICML 2023 · 67 citations
- Promises and Pitfalls of Threshold-based Auto-labelingHarit Vishwakarma, Heguang Lin, Frederic Sala, Ramya Korlakai VinayakNeurIPS 2023 · 16 citations
- Embroid: Unsupervised Prediction Smoothing Can Improve Few-Shot ClassificationNeel Guha, Mayee F. Chen, Kush Bhatia, Azalia Mirhoseini et al.NeurIPS 2023 · 6 citations
- Ground Truth Inference for Weakly Supervised Entity MatchingRenzhi Wu, Alexander Bendeck, Xu Chu, Yeye HeSIGMOD 2023 · 4 citations
