Active Assessment of Prediction Services as Accuracy Surface Over Attribute Combinations
Vihari Piratla, Soumen Chakrabarti, Sunita Sarawagi
摘要
Our goal is to evaluate the accuracy of a black-box classification model, not as a single aggregate on a given test data distribution, but as a surface over a large number of combinations of attributes characterizing multiple test data distributions. Such attributed accuracy measures become important as machine learning models get deployed as a service, where the training data distribution is hidden from clients, and different clients may be interested in diverse regions of the data distribution. We present Attributed Accuracy Assay (AAA) -a Gaussian Process (GP)-based probabilistic estimator for such an accuracy surface. Each attribute combination, called an 'arm', is associated with a Beta density from which the service's accuracy is sampled. We expect the GP to smooth the parameters of the Beta density over related arms to mitigate sparsity. We show that obvious application of GPs cannot address the challenge of heteroscedastic uncertainty over a huge attribute space that is sparsely and unevenly populated. In response, we present two enhancements: pooling sparse observations, and regularizing the scale parameter of the Beta densities. After introducing these innovations, we establish the effectiveness of AAA in terms of both its estimation accuracy and exploration efficiency, through extensive experiments and analysis. Our code and dataset can be found at: https://github.com/vihari/AAA .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper3
- WILDS: A Benchmark of in-the-Wild Distribution ShiftsPang Wei Koh, Shiori Sagawa, Henrik Marklund, Sang Michael Xie 等ICML 2021 · 被引用 1,773 次
- Can I Trust My Fairness Metric? Assessing Fairness with Unlabeled Data and Bayesian InferenceDisi Ji, Padhraic Smyth, Mark SteyversNeurIPS 2020 · 被引用 57 次
- A Programmatic and Semantic Approach to Explaining and Debugging Neural Network Based Object DetectorsEdward Kim, Divya Gopinath, Corina S. Pasareanu, Sanjit A. SeshiaCVPR 2020
相关 Paper
- Probability Distribution of Hypervolume Improvement in Bi-objective Bayesian OptimizationHao Wang, Kaifeng Yang, Michael AffenzellerICML 2024 · 被引用 3 次
- Gaussian Process Probes (GPP) for Uncertainty-Aware ProbingZi Wang, Alexander Ku, Jason Baldridge, Tom Griffiths 等NeurIPS 2023 · 被引用 17 次
- Scalable Variational Bayesian Kernel Selection for Sparse Gaussian Process RegressionTong Teng, Jie Chen, Yehong Zhang, Bryan Kian Hsiang LowAAAI 2020 · 被引用 24 次
- Deep Variational Implicit ProcessesLuis A. Ortega, Simón Rodríguez Santana, Daniel Hernández-LobatoICLR 2023 · 被引用 15 次
- Variational Linearized Laplace Approximation for Bayesian Deep LearningLuis A. Ortega Andrés, Simón Rodríguez Santana, Daniel Hernández-LobatoICML 2024 · 被引用 12 次
