LANCET: Labeling Complex Data at Scale
Huayi Zhang, Lei Cao, Samuel Madden, Elke A. Rundensteiner
摘要
Cutting-edge machine learning techniques often require millions of labeled data objects to train a robust model. Because relying on humans to supply such a huge number of labels is rarely practical, automated methods for label generation are needed. Unfortunately, critical challenges in auto-labeling remain unsolved, including the following research questions: (1) which objects to ask humans to label, (2) how to automatically propagate labels to other objects, and (3) when to stop labeling. These three questions are not only each challenging in their own right, but they also correspond to tightly interdependent problems. Yet existing techniques provide at best isolated solutions to a subset of these challenges. In this work, we propose the first approach, called LANCET, that successfully addresses all three challenges in an integrated framework. LANCET is based on a theoretical foundation characterizing the properties that the labeled dataset must satisfy to train an effective prediction model, namely the Covariate-shift and the Continuity conditions. First, guided by the Covariate-shift condition, LANCET maps raw input data into a semantic feature space, where an unlabeled object is expected to share the same label with its near-by labeled neighbor. Next, guided by the Continuity condition, LANCET selects objects for labeling, aiming to ensure that unlabeled objects always have some sufficiently close labeled neighbors. These two strategies jointly maximize the accuracy of the automatically produced labels and the prediction accuracy of the machine learning models trained on these labels. Lastly, LANCET uses a distribution matching network to verify whether both the Covariate-shift and Continuity conditions hold, in which case it would be safe to terminate the labeling process. Our experiments on diverse public data sets demonstrate that LANCET consistently outperforms the state-of-the-art methods from Snuba to GOGGLES and other baselines by a large margin - up to 30 percentage points increase in accuracy.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- MetaStore: Analyzing Deep Learning Meta-Data at ScaleHuayi Zhang, Binwei Yan, Lei Cao, Samuel Madden 等VLDB 2024 · 被引用 10 次
- VOCALExplore: Pay-as-You-Go Video Data Exploration and Model BuildingMaureen Daum, Enhao Zhang, Dong He, Stephen Mussmann 等VLDB 2023 · 被引用 7 次
- CLaDMoP: Learning Transferrable Models from Successful Clinical Trials via LLMsYiqing Zhang, Xiaozhong Liu, Fabricio MuraiKDD 2025 · 被引用 1 次
- CLID-MU: Cross-Layer Information Divergence Based Meta Update Strategy for Learning with Noisy LabelsRuofan Hu, Dongyu Zhang, Huayi Zhang, Elke A. RundensteinerKDD 2025
它引用的顶会 Paper4
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 被引用 24,064 次
- FixMatch: Simplifying Semi-Supervised Learning with Consistency and ConfidenceKihyuk Sohn, David Berthelot, Nicholas Carlini, Zizhao Zhang 等NeurIPS 2020 · 被引用 5,129 次
- Variational Adversarial Active LearningSamarth Sinha, Sayna Ebrahimi, Trevor DarrellICCV 2019 · 被引用 662 次
- GOGGLES: Automatic Image Labeling with Affinity CodingNilaksh Das, Sanya Chaba, Renzhi Wu, Sakshi Gandhi 等SIGMOD 2020 · 被引用 22 次
相关 Paper
- Pearls from Pebbles: Improved Confidence Functions for Auto-labelingHarit Vishwakarma, Yi Chen, Sui Jiet Tay, Satya Sai Srinath Namburi 等NeurIPS 2024 · 被引用 7 次
- Feed Two Birds with One Scone: Exploiting Wild Data for Both Out-of-Distribution Generalization and DetectionHaoyue Bai, Gregory Canal, Xuefeng Du, Jeongyeol Kwon 等ICML 2023 · 被引用 67 次
- Promises and Pitfalls of Threshold-based Auto-labelingHarit Vishwakarma, Heguang Lin, Frederic Sala, Ramya Korlakai VinayakNeurIPS 2023 · 被引用 16 次
- Embroid: Unsupervised Prediction Smoothing Can Improve Few-Shot ClassificationNeel Guha, Mayee F. Chen, Kush Bhatia, Azalia Mirhoseini 等NeurIPS 2023 · 被引用 6 次
- Ground Truth Inference for Weakly Supervised Entity MatchingRenzhi Wu, Alexander Bendeck, Xu Chu, Yeye HeSIGMOD 2023 · 被引用 4 次
