A Geometric Approach to Predicting Bounds of Downstream Model Performance
Brian J. Goode, Debanjan Datta
Abstract
This paper presents the motivation and methodology for including model application criteria into baseline analysis. We will focus on detailing the interplay between the common measures of mean square error (MSE) and accuracy as it relates to perceived model performance. MSE is a common aggregate measure for the performance of predictive regression models. The advantages are numerous. MSE is agnostic to the choice of model given that the set of possible outcome values are defined on the appropriate metric space. In practice, decisions on how to subsequently use a trained model are based on predictive performance, relative to a baseline where input features are not used - colloquially a "random model". However, the relative performance gains of a model in terms of MSE to the baseline does not guarantee commensurate gains when deployed in downstream applications, systems, or processes. This paper demonstrates one derivation of a distribution to qualify MSE performance for multi-class decision making systems desiring a certain level of accuracy. The model error is qualified through comparison to relevant baselines tied to the application suited to evaluating individual outcome performance criteria.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Related papers
- Expectations vs. Realities: The Cost of MSE-Optimal Forecasting Under Conditional UncertaintyRiku Green, Zahraa S. Abdallah, Telmo de Menezes e Silva FilhoKDD 2026 · 3 citations
- Quantile Risk Control: A Flexible Framework for Bounding the Probability of High-Loss PredictionsJake Snell, Thomas P. Zollo, Zhun Deng, Toniann Pitassi et al.ICLR 2023 · 1 citation
- Rethinking Reward Model Evaluation: Are We Barking up the Wrong Tree?Xueru Wen, Jie Lou, Yaojie Lu, Hongyu Lin et al.ICLR 2025
- When Confidence Meets Accuracy: Exploring the Effects of Multiple Performance Indicators on Trust in Machine Learning ModelsAmy Rechkemmer, Ming YinCHI 2022 · 94 citations
- Needles in the Haystack: Addressing Signal Dilution Improves scRNA-seq Perturbation Response Modeling and EvaluationGabriel Mejia, Henry Miller, Francis Leblanc, BO WANG et al.ICML 2026
