Beyond Model Ranking: Predictability-Aligned Evaluation for Time Series Forecasting
Wanjin Feng, Yuan Yuan, Jingtao Ding, Yong Li
Abstract
In the era of increasingly complex AI models for time series forecasting, progress is often measured by marginal improvements on benchmark leaderboards. However, standard evaluations rely on aggregate metrics (e.g., MSE) that conflate model capability with the intrinsic difficulty of the evaluated instances. To address this, we propose a diagnostic framework anchored in Spectral Coherence Predictability (SCP), which provides an efficient per-instance difficulty reference and yields a corresponding linear MSE lower bound. Complementing this, we introduce the Linear Utilization Ratio (LUR) to quantify how effectively models exploit linearly predictable structures across frequencies. Experiments on synthetic and real-world benchmarks show that SCP aligns strongly with realized forecasting errors across diverse state-of-the-art forecasters. Using this lens, we uncover ``predictability drift,'' revealing that task difficulty is not static but fluctuates significantly over time and variables. Furthermore, stratified evaluation exposes complementary architectural strengths across distinct frequency bands and difficulty regimes. Overall, we advocate moving beyond leaderboard-style ranking toward a more insightful, predictability-aware evaluation that fosters fairer model comparisons and a deeper understanding of model behavior. Code and data are available at https://github.com/WanjinVon/TS_Predictability.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext c6fdbeb2-7714-4217-ba92-24eec04bfc00Builds on6
- Are Transformers Effective for Time Series Forecasting?Ailing Zeng, Muxi Chen, Lei Zhang, Qiang XuAAAI 2023 · 3,619 citations
- FEDformer: Frequency Enhanced Decomposed Transformer for Long-term Series ForecastingTian Zhou, Ziqing Ma, Qingsong Wen, Xue Wang et al.ICML 2022 · 2,912 citations
- A Time Series is Worth 64 Words: Long-term Forecasting with TransformersYuqi Nie, Nam H. Nguyen, Phanwadee Sinthong, Jayant KalagnanamICLR 2023 · 536 citations
- TimesNet: Temporal 2D-Variation Modeling for General Time Series AnalysisHaixu Wu, Tengge Hu, Yong Liu, Hang Zhou et al.ICLR 2023 · 423 citations
- Quantifying and Estimating the Predictability Upper Bound of Univariate Numeric Time SeriesJamal Mohammed, Michael H. Böhlen, Sven HelmerKDD 2024 · 1 citation
Related papers
- Time-Series Decomposition as a Standalone Task: A Mechanism-Driven Diagnostic BenchmarkZipeng Wu, Jiani Wei, Shiqiao Zhou, Jiajun Chen et al.ICML 2026
- TSFAdv: Frequency-Guided Black-Box Adversarial Attacks on Time Series ForecastingQizhuo Han, Xiangrui Cai, Sihan Xu, Ying Zhang et al.ICML 2026
- It's TIME: Towards the Next Generation of Time Series Forecasting BenchmarksZhongzheng Qiao, SHENG PAN, Anni Wang, Viktoriya Zhukova et al.ICML 2026 · 9 citations
- From Observations to States: Latent Time Series ForecastingJie Yang, Yifan Hu, Yuante Li, Kexin Zhang et al.ICML 2026 · 3 citations
- How Robust are Model Rankings : A Leaderboard Customization Approach for Equitable EvaluationSwaroop Mishra, Anjana ArunkumarAAAI 2021 · 27 citations
