Evaluating Large Language Models on Time Series Feature Understanding: A Comprehensive Taxonomy and Benchmark
Elizabeth Fons, Rachneet Kaur, Soham Palande, Zhen Zeng, Tucker Balch, Manuela Veloso, Svitlana Vyetrenko
摘要
Large Language Models (LLMs) offer the potential for automatic time series analysis and reporting, which is a critical task across many domains, spanning healthcare, finance, climate, energy, and many more. In this paper, we propose a framework for rigorously evaluating the capabilities of LLMs on time series understanding, encompassing both univariate and multivariate forms. We introduce a comprehensive taxonomy of time series features, a critical framework that delineates various characteristics inherent in time series data. Leveraging this taxonomy, we have systematically designed and synthesized a diverse dataset of time series, embodying the different outlined features, each accompanied by textual descriptions. This dataset acts as a solid foundation for assessing the proficiency of LLMs in comprehending time series. Our experiments shed light on the strengths and limitations of stateof-the-art LLMs in time series understanding, revealing which features these models readily comprehend effectively and where they falter. In addition, we uncover the sensitivity of LLMs to factors including the formatting of the data, the position of points queried within a series and the overall time series length.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper8
- ChatTime: A Unified Multimodal Time Series Foundation Model Bridging Numerical and Textual DataChengsen Wang, Qi Qi, Jingyu Wang, Haifeng Sun 等AAAI 2025 · 被引用 109 次
- ChatTS: Aligning Time Series with LLMs via Synthetic Data for Enhanced Understanding and ReasoningZhe Xie, Zeyan Li, Xiao He, Longlong Xu 等VLDB 2025 · 被引用 87 次
- BEDTime: A Unified Benchmark for Automatically Describing Time SeriesMedhasweta Sen, Zachary Gottesman, Jiaxing Qiu, C. Bayan Bruss 等ICML 2026 · 被引用 8 次
- ReTabAD: A Benchmark for Restoring Semantic Context in Tabular Anomaly DetectionSanghyu Yoon, Dongmin Kim, Suhee Yoon, Ye Seul Sim 等ICLR 2026 · 被引用 3 次
- Peak-Detector: Explainable Peak Detection via Instruction-Tuned Large Language Models in Physiological SignalJiahui Li, Yida Zhang, Zixuan Zeng, Jiayu Chen 等UbiComp 2026 · 被引用 1 次
它引用的顶会 Paper5
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Calibrate Before Use: Improving Few-shot Performance of Language ModelsZihao Zhao, Eric Wallace, Shi Feng, Dan Klein 等ICML 2021 · 被引用 1,843 次
- One Fits All: Power General Time Series Analysis by Pretrained LMTian Zhou, Peisong Niu, Xue Wang, Liang Sun 等NeurIPS 2023 · 被引用 1,178 次
- Large Language Models Are Zero-Shot Time Series ForecastersNate Gruver, Marc Finzi, Shikai Qiu, Andrew Gordon WilsonNeurIPS 2023 · 被引用 898 次
- Llemma: An Open Language Model for MathematicsZhangir Azerbayev, Hailey Schoelkopf, Keiran Paster, Marco Dos Santos 等ICLR 2024 · 被引用 433 次
相关 Paper
- TimeSeriesExamAgent: Creating Time Series Reasoning Benchmarks at ScaleMalgorzata Gwiazda, Yifu Cai, Mononito Goswami, Arjun Choudhry 等ICLR 2026 · 被引用 6 次
- Understanding Why Large Language Models Can Be Ineffective in Time Series Analysis: The Impact of Modality AlignmentLiangwei Nathan Zheng, Chang George Dong, Wei Emma Zhang, Lin Yue 等KDD 2025 · 被引用 1 次
- TsLLM: Augmenting LLMs for General Time Series Understanding and PredictionFelix Parker, Nimeesha Chan, Chi Zhang, Kimia GhobadiICML 2026 · 被引用 3 次
- TimeCAP: Learning to Contextualize, Augment, and Predict Time Series Events with Large Language Model AgentsGeon Lee, Wenchao Yu, Kijung Shin, Wei Cheng 等AAAI 2025 · 被引用 39 次
- Time-MQA: Time Series Multi-Task Question Answering with Context EnhancementYaxuan Kong, Yiyuan Yang, Yoontae Hwang, Wenjie Du 等ACL 2025
