Active Assessment of Prediction Services as Accuracy Surface Over Attribute Combinations
Vihari Piratla, Soumen Chakrabarti, Sunita Sarawagi
Abstract
Our goal is to evaluate the accuracy of a black-box classification model, not as a single aggregate on a given test data distribution, but as a surface over a large number of combinations of attributes characterizing multiple test data distributions. Such attributed accuracy measures become important as machine learning models get deployed as a service, where the training data distribution is hidden from clients, and different clients may be interested in diverse regions of the data distribution. We present Attributed Accuracy Assay (AAA) -a Gaussian Process (GP)-based probabilistic estimator for such an accuracy surface. Each attribute combination, called an 'arm', is associated with a Beta density from which the service's accuracy is sampled. We expect the GP to smooth the parameters of the Beta density over related arms to mitigate sparsity. We show that obvious application of GPs cannot address the challenge of heteroscedastic uncertainty over a huge attribute space that is sparsely and unevenly populated. In response, we present two enhancements: pooling sparse observations, and regularizing the scale parameter of the Beta densities. After introducing these innovations, we establish the effectiveness of AAA in terms of both its estimation accuracy and exploration efficiency, through extensive experiments and analysis. Our code and dataset can be found at: https://github.com/vihari/AAA .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 0f11d94c-7f53-4777-a46d-56aa2d30f880Cited by top-tier papers1
Ask how each one uses itBuilds on3
- WILDS: A Benchmark of in-the-Wild Distribution ShiftsPang Wei Koh, Shiori Sagawa, Henrik Marklund, Sang Michael Xie et al.ICML 2021 · 1,773 citations
- Can I Trust My Fairness Metric? Assessing Fairness with Unlabeled Data and Bayesian InferenceDisi Ji, Padhraic Smyth, Mark SteyversNeurIPS 2020 · 57 citations
- A Programmatic and Semantic Approach to Explaining and Debugging Neural Network Based Object DetectorsEdward Kim, Divya Gopinath, Corina S. Pasareanu, Sanjit A. SeshiaCVPR 2020
Related papers
- Probability Distribution of Hypervolume Improvement in Bi-objective Bayesian OptimizationHao Wang, Kaifeng Yang, Michael AffenzellerICML 2024 · 3 citations
- Gaussian Process Probes (GPP) for Uncertainty-Aware ProbingZi Wang, Alexander Ku, Jason Baldridge, Tom Griffiths et al.NeurIPS 2023 · 17 citations
- Scalable Variational Bayesian Kernel Selection for Sparse Gaussian Process RegressionTong Teng, Jie Chen, Yehong Zhang, Bryan Kian Hsiang LowAAAI 2020 · 24 citations
- Deep Variational Implicit ProcessesLuis A. Ortega, Simón Rodríguez Santana, Daniel Hernández-LobatoICLR 2023 · 15 citations
- Variational Linearized Laplace Approximation for Bayesian Deep LearningLuis A. Ortega Andrés, Simón Rodríguez Santana, Daniel Hernández-LobatoICML 2024 · 12 citations
