TimeOmni-VL: Unified Models for Time Series Understanding and Generation
Tong Guan, SHENG PAN, Johan Barthelemy, Zhao Li, Yujun Cai, Cesare Alippi, Ming Jin, Shirui Pan
Abstract
Recent time series modeling faces a sharp divide between numerical generation and semantic understanding, with research showing that generation models often rely on superficial pattern matching, while understanding-oriented models struggle with high-fidelity numerical output. Although unified multimodal models (UMMs) have bridged this gap in vision, their potential for time series remains untapped. We propose TIMEOMNI-VL, the first vision-centric framework that unifies time series understanding and generation through two key innovations: (1) Fidelity-preserving bidirectional mapping between time series and images (Bi-TSI), which advances Time Series-to-Image (TS2I) and Image-to-Time Series (I2TS) conversions to ensure near-lossless transformations. (2) Understanding-guided generation. We introduce TSUMM-SUITE, a novel dataset consisting of six understanding tasks rooted in time series analytics and coupled with two generation tasks. With a calibrated Chain-of-Thought (CoT), TIMEOMNI-VL is the first to leverage time series understanding as an explicit control signal for high-fidelity generation. Experiments confirm that this unified approach significantly improves semantic understanding and numerical precision, establishing a new frontier for multimodal time series modeling. Code 1 and checkpoint 2 are publicly available.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on17
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- GLIDE: Towards Photorealistic Image Generation and Editing with Text-Guided Diffusion ModelsAlexander Quinn Nichol, Prafulla Dhariwal, Aditya Ramesh, Pranav Shyam et al.ICML 2022 · 4,691 citations
- Connecting the Dots: Multivariate Time Series Forecasting with Graph Neural NetworksZonghan Wu, Shirui Pan, Guodong Long, Jing Jiang et al.KDD 2020 · 1,738 citations
- Time-LLM: Time Series Forecasting by Reprogramming Large Language ModelsMing Jin, Shiyu Wang, Lintao Ma, Zhixuan Chu et al.ICLR 2024 · 915 citations
- Unified Training of Universal Time Series Forecasting TransformersGerald Woo, Chenghao Liu, Akshat Kumar, Caiming Xiong et al.ICML 2024 · 513 citations
Related papers
- TimeOmni-1: Incentivizing Complex Reasoning with Time Series in Large Language ModelsTong Guan, Zijie Meng, Dianqi Li, Shiyu Wang et al.ICLR 2026 · 29 citations
- UnifiedVisual: A Framework for Constructing Unified Vision-Language DatasetsPengyu Wang, Shaojun Zhou, Chenkun Tan, Xinghao Wang et al.EMNLP 2025
- SciTS: Scientific Time Series Understanding and Generation with LLMsWen Wu, Ziyang Zhang, Liwei Liu, Xuenan Xu et al.ICLR 2026 · 11 citations
- Time-VLM: Exploring Multimodal Vision-Language Models for Augmented Time Series ForecastingSiru Zhong, Weilin Ruan, Ming Jin, Huan Li et al.ICML 2025
- Omni-View: Unlocking How Generation Facilitates Understanding in Unified 3D Model based on Multiview imagesJiaKui Hu, Shanshan Zhao, Qing-Guo Chen, Xuerui Qiu et al.ICLR 2026 · 20 citations
