Evaluating State-of-the-Art Classification Models Against Bayes Optimality
Ryan Theisen, Huan Wang, Lav R. Varshney, Caiming Xiong, Richard Socher
摘要
Evaluating the inherent difficulty of a given data-driven classification problem is important for establishing absolute benchmarks and evaluating progress in the field. To this end, a natural quantity to consider is the Bayes error, which measures the optimal classification error theoretically achievable for a given data distribution. While generally an intractable quantity, we show that we can compute the exact Bayes error of generative models learned using normalizing flows. Our technique relies on a fundamental result, which states that the Bayes error is invariant under invertible transformation. Therefore, we can compute the exact Bayes error of the learned flow models by computing it for Gaussian base distributions, which can be done efficiently using Holmes-Diaconis-Ross integration. Moreover, we show that by varying the temperature of the learned flow models, we can generate synthetic datasets that closely resemble standard benchmark datasets, but with almost any desired Bayes error. We use our approach to conduct a thorough investigation of state-of-the-art classification models, and find that in some -- but not all -- cases, these models are capable of obtaining accuracy very near optimal. Finally, we use our method to evaluate the intrinsic"hardness"of standard benchmark datasets, and classes within those datasets.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper7
- Demystifying the Optimal Performance of Multi-Class ClassificationMinoh Jeong, Martina Cardone, Alex DytsoNeurIPS 2023 · 被引用 17 次
- Practical estimation of the optimal classification error with soft labels and calibrationRyota Ushio, Takashi Ishida, Masashi SugiyamaICLR 2026 · 被引用 6 次
- Certified Robust Accuracy of Neural Networks Are Bounded Due to Bayes ErrorsRuihan Zhang, Jun SunCAV 2024 · 被引用 5 次
- CapBencher: Give Your LLM Benchmark a Built-in Alarm for Test-Set OverfittingTakashi Ishida, Thanawat Lodkaew, Ikko YamaneICML 2026 · 被引用 4 次
- Is the Performance of My Deep Network Too Good to Be True? A Direct Approach to Estimating the Bayes Error in Binary ClassificationTakashi Ishida, Ikko Yamane, Nontawat Charoenphakdee, Gang Niu 等ICLR 2023 · 被引用 3 次
它引用的顶会 Paper4
- Random Erasing Data AugmentationZhun Zhong, Liang Zheng, Guoliang Kang, Shaozi Li 等AAAI 2020 · 被引用 4,134 次
- Sharpness-aware Minimization for Efficiently Improving GeneralizationPierre Foret, Ariel Kleiner, Hossein Mobahi, Behnam NeyshaburICLR 2021 · 被引用 1,861 次
- Understanding the Limitations of Conditional Generative ModelsEthan Fetaya, Jörn-Henrik Jacobsen, Will Grathwohl, Richard S. ZemelICLR 2020 · 被引用 65 次
- Efficient Neural Vision Systems Based on Convolutional Image AcquisitionPedram Pad, Simon Narduzzi, Clément Kündig, Engin Türetken 等CVPR 2020
相关 Paper
- Training Normalizing Flows with the Information Bottleneck for Competitive Generative ClassificationLynton Ardizzone, Radek Mackowiak, Carsten Rother, Ullrich KötheNeurIPS 2020 · 被引用 62 次
- Posterior Refinement Improves Sample Efficiency in Bayesian Neural NetworksAgustinus Kristiadi, Runa Eschenhagen, Philipp HennigNeurIPS 2022 · 被引用 17 次
- Inflationary Flows: Calibrated Bayesian Inference with Diffusion-Based ModelsDaniela de Albuquerque, John M. PearsonNeurIPS 2024 · 被引用 3 次
- Composing Normalizing Flows for Inverse ProblemsJay Whang, Erik M. Lindgren, Alex DimakisICML 2021 · 被引用 56 次
- FALCON: Few-step Accurate Likelihoods for Continuous FlowsDanyal Rehman, Tara Akhound-Sadegh, Artem Gazizov, Yoshua Bengio 等ICLR 2026 · 被引用 13 次
