Harnessing Vision-Language Models for Time Series Anomaly Detection
Zelin He, Sarah Alnegheimish, Matthew Reimherr
Abstract
Time-series anomaly detection (TSAD) has played a vital role in a variety of fields, including healthcare, finance, and sensor-based condition monitoring. Prior methods, which mainly focus on training domain-specific models on numerical data, lack the visual-temporal understanding capacity that human experts have to identify contextual anomalies. To fill this gap, we explore a solution based on vision language models (VLMs). Recent studies have shown the ability of VLMs for visual understanding tasks, yet their direct application to time series has fallen short on both accuracy and efficiency. To harness the power of VLMs for TSAD, we propose a two-stage solution, with (1) ViT4TS, a vision-screening stage built on a relatively lightweight pre-trained vision encoder, which leverages 2-D time series representations to accurately localize candidate anomalies; (2) VLM4TS, a VLM-based stage that integrates global temporal context and VLM's visual understanding capacity to refine the detection upon the candidates provided by ViT4TS. We show that without any time-series training, VLM4TS outperforms time-series pretrained and from-scratch baselines in most cases, yielding a 24.6% improvement in F1-max score over the best baseline. Moreover, VLM4TS also consistently outperforms existing language model-based TSAD methods and is on average 36 × more efficient in token usage.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d8d76948-1c28-4f14-a252-5de4bce1092eCited by top-tier papers2
- Detection of unknown unknowns in autonomous systemsAyan Banerjee, Sandeep GuptaICLR 2026
- AnomSeer: Reinforcing Multimodal LLMs to Reason for Time-Series Anomaly DetectionJunru Zhang, Lang Feng, Haoran Shi, Xu Guo et al.ICML 2026
Builds on10
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
- Anomaly Transformer: Time Series Anomaly Detection with Association DiscrepancyJiehui Xu, Haixu Wu, Jianmin Wang, Mingsheng LongICLR 2022 · 960 citations
- A decoder-only foundation model for time-series forecastingAbhimanyu Das, Weihao Kong, Rajat Sen, Yichen ZhouICML 2024 · 601 citations
- UniTS: A Unified Multi-Task Time Series ModelShanghua Gao, Teddy Koker, Owen Queen, Tom Hartvigsen et al.NeurIPS 2024 · 159 citations
- Time Series as Images: Vision Transformer for Irregularly Sampled Time SeriesZekun Li, Shiyang Li, Xifeng YanNeurIPS 2023 · 145 citations
Related papers
- ViTs: Teaching Machines to See Time Series Anomalies Like Human ExpertsZexin Wang, Changhua Pei, Yang Liu, Hengyue Jiang et al.WWW 2026
- Can Multimodal LLMs Perform Time Series Anomaly Detection?Xiongxiao Xu, Haoran Wang, Yueqing Liang, Philip S. Yu et al.WWW 2026 · 18 citations
- Can LLMs Understand Time Series Anomalies?Zihao Zhou, Rose YuICLR 2025
- TsLLM: Augmenting LLMs for General Time Series Understanding and PredictionFelix Parker, Nimeesha Chan, Chi Zhang, Kimia GhobadiICML 2026 · 3 citations
- Delving into Large Language Models for Effective Time-Series Anomaly DetectionJunwoo Park, Kyudan Jung, Dohyun Lee, Hyuck Lee et al.NeurIPS 2025 · 1 citation
