Expectations vs. Realities: The Cost of MSE-Optimal Forecasting Under Conditional Uncertainty
Riku Green, Zahraa S. Abdallah, Telmo de Menezes e Silva Filho
Abstract
Multi-step time series forecasting (MSF) is commonly evaluated using point-wise error metrics such as mean squared error (MSE), implicitly treating the conditional mean as a sufficient target. We show that this can be misleading under conditional uncertainty, where the conditional expectation becomes unrepresentative of typical realized values at longer horizons. We formalize this effect through a conditional uncertainty gap and prove that whenever this gap is nonzero, no deterministic predictor can simultaneously minimize MSE and match the marginal distribution of realized futures. This establishes a fundamental, model-agnostic trade-off between point accuracy and marginal realism in MSF evaluation. Using controlled stochastic dynamical systems and nine real-world forecasting benchmarks, we empirically characterize the resulting accuracy--realism frontier and quantify the practical cost of MSE-only model selection. As conditional uncertainty increases with forecast horizon, the attainable set expands into a pronounced Pareto front, separating MSE-optimal but under-dispersed predictors from methods that trade accuracy for realistic marginal variability. Across benchmarks, we find that small relaxations in MSE (≤ 5%) frequently unlock disproportionate gains in marginal realism, with median improvements of 17.3% and gains exceeding 30% in some datasets. We further show that common forecasting strategies systematically occupy different regions of this frontier: direct multi-output predictors concentrate near the accuracy-optimal extreme, while recursive strategies and sample-based inference favors marginal realism. Together, these results expose a structural failure mode of MSE-based evaluation in long-horizon forecasting and recast strategy and inference selection as navigation of an unavoidable accuracy--realism trade-off.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on4
- TFB: Towards Comprehensive and Fair Benchmarking of Time Series Forecasting MethodsXiangfei Qiu, Jilin Hu, Lekui Zhou, Xingjian Wu et al.VLDB 2024 · 292 citations
- Learning the Pareto Front with HypernetworksAviv Navon, Aviv Shamsian, Ethan Fetaya, Gal ChechikICLR 2021 · 189 citations
- Probabilistic Time Series Forecasting with Shape and Temporal DiversityVincent Le Guen, Nicolas ThomeNeurIPS 2020 · 34 citations
- AutoXPCR: Automated Multi-Objective Model Selection for Time Series ForecastingRaphael Fischer, Amal SaadallahKDD 2024 · 7 citations
Related papers
- Regions of Reliability in the Evaluation of Multivariate Probabilistic ForecastsÉtienne Marcotte, Valentina Zantedeschi, Alexandre Drouin, Nicolas ChapadosICML 2023 · 10 citations
- A Geometric Approach to Predicting Bounds of Downstream Model PerformanceBrian J. Goode, Debanjan DattaKDD 2020
- DistDF: Time-series Forecasting Needs Joint-distribution Wasserstein AlignmentEric Wang, Licheng Pan, Yuan Lu, Zhixuan Chu et al.ICLR 2026 · 19 citations
- Conformal Time-series ForecastingKamile Stankeviciute, Ahmed M. Alaa, Mihaela van der SchaarNeurIPS 2021 · 233 citations
- What if Tomorrow is the World Cup Final? Counterfactual Time Series Forecasting with Textual ConditionsShuqi Gu, Yongxiang Zhao, Baoyu Jing, Kan RenICML 2026
