Finding Label and Model Errors in Perception Data With Learned Observation Assertions
Daniel Kang, Nikos Aréchiga, Sudeep Pillai, Peter D. Bailis, Matei Zaharia
摘要
ML is being deployed in complex, real-world scenarios where errors have impactful consequences. In these systems, thorough testing of the ML pipelines is critical. A key component in ML deployment pipelines is the curation of labeled training data. Common practice in the ML literature assumes that labels are the ground truth. However, in our experience in a large autonomous vehicle development center, we have found that vendors can often provide erroneous labels, which can lead to downstream safety risks in trained models.
To address these issues, we propose a new abstraction, learned observation assertions, and implement it in a system called Fixy. Fixy leverages existing organizational resources, such as existing (possibly noisy) labeled datasets or previously trained ML models, to learn a probabilistic model for finding errors in human-or modelgenerated labels. Given user-provided features and these existing resources, Fixy learns feature distributions that specify likely and unlikely values (e.g., that a speed of 30mph is likely but 300mph is unlikely). It then uses these feature distributions to score labels for potential errors. We show that Fixy can automatically rank potential errors in real datasets with up to 2× higher precision compared to recent work on model assertions and standard techniques such as uncertainty sampling.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它相关 Paper
- Towards Observability for Production Machine Learning Pipelines [Vision]Shreya Shankar, Aditya G. ParameswaranVLDB 2022 · 被引用 21 次
- Auto-Validate: Unsupervised Data Validation Using Data-Domain Patterns Inferred from Data LakesJie Song, Yeye HeSIGMOD 2021 · 被引用 26 次
- FILA: Online Auditing of Machine Learning Model Accuracy under Finite Labelling BudgetNaiqing Guan, Nick KoudasSIGMOD 2022 · 被引用 1 次
- Deep k-NN for Noisy LabelsDara Bahri, Heinrich Jiang, Maya R. GuptaICML 2020 · 被引用 90 次
- Logs In, Patches Out: Automated Vulnerability Repair via Tree-of-Thought LLM AnalysisYoungjoon Kim, Sunguk Shin, Hyoungshick Kim, Jiwon YoonUSENIX Security 2025
