Beyond Model Ranking: Predictability-Aligned Evaluation for Time Series Forecasting
Wanjin Feng, Yuan Yuan, Jingtao Ding, Yong Li
摘要
In the era of increasingly complex AI models for time series forecasting, progress is often measured by marginal improvements on benchmark leaderboards. However, standard evaluations rely on aggregate metrics (e.g., MSE) that conflate model capability with the intrinsic difficulty of the evaluated instances. To address this, we propose a diagnostic framework anchored in Spectral Coherence Predictability (SCP), which provides an efficient per-instance difficulty reference and yields a corresponding linear MSE lower bound. Complementing this, we introduce the Linear Utilization Ratio (LUR) to quantify how effectively models exploit linearly predictable structures across frequencies. Experiments on synthetic and real-world benchmarks show that SCP aligns strongly with realized forecasting errors across diverse state-of-the-art forecasters. Using this lens, we uncover ``predictability drift,'' revealing that task difficulty is not static but fluctuates significantly over time and variables. Furthermore, stratified evaluation exposes complementary architectural strengths across distinct frequency bands and difficulty regimes. Overall, we advocate moving beyond leaderboard-style ranking toward a more insightful, predictability-aware evaluation that fosters fairer model comparisons and a deeper understanding of model behavior. Code and data are available at https://github.com/WanjinVon/TS_Predictability.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper6
- Are Transformers Effective for Time Series Forecasting?Ailing Zeng, Muxi Chen, Lei Zhang, Qiang XuAAAI 2023 · 被引用 3,619 次
- FEDformer: Frequency Enhanced Decomposed Transformer for Long-term Series ForecastingTian Zhou, Ziqing Ma, Qingsong Wen, Xue Wang 等ICML 2022 · 被引用 2,912 次
- A Time Series is Worth 64 Words: Long-term Forecasting with TransformersYuqi Nie, Nam H. Nguyen, Phanwadee Sinthong, Jayant KalagnanamICLR 2023 · 被引用 536 次
- TimesNet: Temporal 2D-Variation Modeling for General Time Series AnalysisHaixu Wu, Tengge Hu, Yong Liu, Hang Zhou 等ICLR 2023 · 被引用 423 次
- Quantifying and Estimating the Predictability Upper Bound of Univariate Numeric Time SeriesJamal Mohammed, Michael H. Böhlen, Sven HelmerKDD 2024 · 被引用 1 次
相关 Paper
- Time-Series Decomposition as a Standalone Task: A Mechanism-Driven Diagnostic BenchmarkZipeng Wu, Jiani Wei, Shiqiao Zhou, Jiajun Chen 等ICML 2026
- TSFAdv: Frequency-Guided Black-Box Adversarial Attacks on Time Series ForecastingQizhuo Han, Xiangrui Cai, Sihan Xu, Ying Zhang 等ICML 2026
- It's TIME: Towards the Next Generation of Time Series Forecasting BenchmarksZhongzheng Qiao, SHENG PAN, Anni Wang, Viktoriya Zhukova 等ICML 2026 · 被引用 9 次
- From Observations to States: Latent Time Series ForecastingJie Yang, Yifan Hu, Yuante Li, Kexin Zhang 等ICML 2026 · 被引用 3 次
- How Robust are Model Rankings : A Leaderboard Customization Approach for Equitable EvaluationSwaroop Mishra, Anjana ArunkumarAAAI 2021 · 被引用 27 次
