From Teacher Pathways to Invariant Manifolds: Consensus Subspace Distillation for TSFMs
Zexing Zhang, Tianyang Lei, Jichao Li, Yang Kewei
摘要
Time-series foundation models (TSFMs) deliver strong cross-domain generalization, but their scale makes deployment costly. Knowledge distillation is a natural compression route, yet prior TSFM distillation typically imitates teacher outputs, features, or pairwise relations, and therefore remains tightly coupled to teacher-specific training trajectories while underutilizing two empirical properties: (i) high-level representations across model scales tend to converge toward a shared, approximately low-rank geometry, and (ii) layer-wise utility follows a long-tail pattern. We propose consensus subspace distillation, which reframes distillation as aligning a student to a model-agnostic geometric object: a scale-invariant low-rank consensus subspace together with its center statistics. Offline, we screen high-contribution layers via drop-layer marginal loss, estimate a shrinkage-stabilized covariance from their embeddings, and derive a truncated eigensubspace that defines a consensus projector. Online, we project student embeddings into this subspace and match the teacher’s projected mean and covariance using a lightweight mean--covariance objective, enabling stable optimization without rigid pointwise feature binding. To mitigate subset-induced bias, we further introduce a frequency-domain uncertainty injection mechanism that inflates spectral density based on characteristic-function discrepancies and injects dispersion only within the consensus directions. Across forecasting and imputation, the distilled student matches or slightly improves upon the teacher, while exhibiting a predictable trade-off under strict zero-shot classification. With MOMENT-Large as teacher, we achieve about 90% parameter reduction and substantial distillation-time savings while retaining comparable performance across multiple time-series tasks. Code and compressed weights are available at anonymous.4open.science/r/CSD-13C3/.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper12
- Contrastive Representation DistillationYonglong Tian, Dilip Krishnan, Phillip IsolaICLR 2020 · 被引用 1,305 次
- A Time Series is Worth 64 Words: Long-term Forecasting with TransformersYuqi Nie, Nam H. Nguyen, Phanwadee Sinthong, Jayant KalagnanamICLR 2023 · 被引用 536 次
- MOMENT: A Family of Open Time-series Foundation ModelsMononito Goswami, Konrad Szafer, Arjun Choudhry, Yifu Cai 等ICML 2024 · 被引用 442 次
- Evaluating Quantized Large Language ModelsShiyao Li, Xuefei Ning, Luning Wang, Tengxuan Liu 等ICML 2024 · 被引用 88 次
- From Dense to Sparse: Contrastive Pruning for Better Pre-trained Language Model CompressionRunxin Xu, Fuli Luo, Chengyu Wang, Baobao Chang 等AAAI 2022 · 被引用 32 次
相关 Paper
- TS-Memory: Plug-and-Play Memory for Time Series Foundation ModelsSisuo Lyu, Siru Zhong, Tiegang Chen, Weilin Ruan 等KDD 2026
- Generalizable Knowledge Distillation from Vision Foundation Models for Semantic SegmentationChonghua Lv, Dong Zhao, Shuang Wang, Dou Quan 等CVPR 2026 · 被引用 1 次
- S^2-KD: Semantic-Spectral Knowledge Distillation Spatiotemporal ForecastingWenshuo Wang, Yaomin Shen, Yingjie Tan, Yihao ChenAAAI 2026 · 被引用 5 次
- Beyond Point Predictions: Manifold Expansion and Dual Alignment for Robust Time Series DistillationJunyao Hong, Zesheng Lai, Xinyi Xiao, Suyang Zhou 等ICML 2026
- Harmonic Dataset Distillation for Time Series ForecastingSeungha Hong, Sanghwan Jang, Wonbin Kweon, Suyeon Kim 等AAAI 2026
