StarEmbed: Benchmarking Time Series Foundation Models on Astronomical Observations of Variable Stars
Weijian Li, Hong-Yu Chen, Nabeel Rehemtulla, Ved Shah, Dongho Kim, Dennis Wu, Qinjie Lin, Adam Miller, Han Liu
Abstract
Current time series foundation model (TSFM) training corpora largely omit data with certain complexities like irregular temporal sampling. Astronomical time series of stellar fluxes (``light curves'') are available in immense quantities and exhibit irregular sampling, multiple variates, and heteroskedasticity. We introduce , the first public benchmark for light curves comprised of real observations of 40,000 stars across seven classes and evaluations in clustering, classification, and out-of-distribution (OOD) source detection. We benchmark TSFMs with differing architecture and training strategies as well as domain-specific transformers. Our results demonstrate that the family, despite being pre-trained on regularly sampled non-astronomical data, yields state-of-the-art (SOTA) performance in light curve clustering and OOD detection. While no TSFM strictly surpasses the classification performance of the long-established domain baseline, they do demonstrate excellent generalization abilities. marks a step toward universal light curve embeddings and improved TSFM performance on challenging data.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 1be0734f-cf9e-4c6d-a497-3f0fe9ac8a34Builds on11
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou et al.ICCV 2021 · 8,921 citations
- Informer: Beyond Efficient Transformer for Long Sequence Time-Series ForecastingHaoyi Zhou, Shanghang Zhang, Jieqi Peng, Shuai Zhang et al.AAAI 2021 · 7,289 citations
- TS2Vec: Towards Universal Representation of Time SeriesZhihan Yue, Yujing Wang, Juanyong Duan, Tianmeng Yang et al.AAAI 2022 · 938 citations
- A Time Series is Worth 64 Words: Long-term Forecasting with TransformersYuqi Nie, Nam H. Nguyen, Phanwadee Sinthong, Jayant KalagnanamICLR 2023 · 536 citations
- Unified Training of Universal Time Series Forecasting TransformersGerald Woo, Chenghao Liu, Akshat Kumar, Caiming Xiong et al.ICML 2024 · 513 citations
Related papers
- TSB-AutoAD: Towards Automated Solutions for Time-Series Anomaly Detection [E, A & B]Qinghua Liu, Seunghak Lee, John PaparrizosVLDB 2025 · 13 citations
- Federated Foundation Models on Heterogeneous Time SeriesShengchao Chen, Guodong Long, Jing Jiang, Chengqi ZhangAAAI 2025 · 5 citations
- Understanding the Implicit Biases of Design Choices for Time Series Foundation ModelsAnnan Yu, Danielle C. Maddix, Boran Han, Xiyuan Zhang et al.ICLR 2026 · 11 citations
- FeDaL: Federated Dataset Learning for General Time Series Foundation ModelsShengchao Chen, Guodong Long, Michael Blumenstein, Jing JiangICLR 2026 · 11 citations
- Sundial: A Family of Highly Capable Time Series Foundation ModelsYong Liu, Guo Qin, Zhiyuan Shi, Zhi Chen et al.ICML 2025
