Lune

NeurIPS2025Top-tier venue

Unified Transferability Metrics for Time Series Foundation Models

Weiyang Zhang, Xinyang Chen, Xiucheng Li, Kehai Chen, Weili Guan, Liqiang Nie

2025Year
5Citations

Abstract

With the increasing number of time series pre-trained models, designing transferability evaluation metrics for time series has become an urgent problem to address. While transferability evaluation has been extensively studied in computer vision, we aim to address a critical gap by developing tailored metrics for time series analysis. In this paper, we introduce TEMPLATE, a transferability estimation framework specifically tailored for versatile time series analysis, comprising three complementary metrics: (1) Dependency Learning Score quantifies a model's capacity to capture temporal dependencies. (2) Pattern Learning Score evaluates the representation quality in extracting discriminative temporal patterns. (3) Task Adaptation Score assesses cross-task generalization capability, enabling versatile time series analysis. TEMPLATE presents a versatile framework compatible with both classification and regression paradigms. Through comprehensive benchmarking across 5 distinct downstream tasks, our method demonstrates superior capability in identifying optimal pre-trained models from heterogeneous model pools for transfer learning. Compared to the state-of-the-art method ETran, our approach improves the weighted Kendall's τ w across 5 downstream tasks by 35%. The code is available at https://github.com/TEMPLATE.

Recently, pre-trained models have drawn increasing attention in the time series domain due to their exceptional performance in computer vision and natural language processing [1]. These models have achieved significant success across various time series downstream tasks [2, 3] and are readily available on platforms like HuggingFace [4] and TensorFlow Hub [5]. However, no single model consistently outperforms others across all datasets. Therefore, selecting the most suitable pre-trained time series model for a given target task has become a pressing challenge. A time-consuming solution is to fine-tune all pre-trained models on the target dataset and then select the best-performing fine-tuned model. But compared to pre-trained models in computer vision such as ResNet [6] and MobileNet [7], time-series pre-trained models have much larger parameter scales [8, 9], making direct brute-force fine-tuning incur enormous time costs and high computational resource requirements [10], as shown in the left part of Figure 1. Recent studies propose fast transferability evaluation methods to efficiently rank models and select the optimal one. Existing methods can generally be categorized into static and dynamic approaches [11]. Static methods calculate scores directly based on the statistical information of the model, such as LEEP [12], NLEEP [13], H-score [14] and TMI [15]. In contrast, dynamic methods transform this statistical information using certain learning frameworks or representation space mapping algorithms before calculating scores, such as SFDA [16], LogME [17], and ETran [18]. These approaches are empirically validated as effective metrics for selecting computer vision models. 39th Conference on Neural Information Processing Systems (NeurIPS 2025).

0 20 40 60 80

Fine-tune (h) Metrics (Acc, MSE, F1-score etc.

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

lune papers fulltext ba189fa8-1932-493a-a5f0-b3a924e0cdce

Builds on23

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines