TSVer: A Benchmark for Fact Verification Against Time-Series Evidence
Marek Strong, Andreas Vlachos
Abstract
Reasoning over temporal and numerical data, such as time series, is a crucial aspect of factchecking. While many systems have recently been developed to handle this form of evidence, their evaluation remains limited by existing datasets, which often lack structured evidence, provide insufficient justifications for verdicts, or rely on synthetic claims. In this paper, we introduce TSVER, a new benchmark dataset for fact verification focusing on temporal and numerical reasoning with time-series evidence. TSVER contains 304 real-world claims sourced from 41 fact-checking organizations and a curated database of 400 time series covering diverse domains. Each claim is annotated with time frames across all pertinent time series, along with a verdict and justifications reflecting how the evidence is used to reach the verdict. Using an LLM-assisted multi-step annotation process, we improve the quality of our annotations and achieve an inter-annotator agreement of κ = 0.77 on verdicts. We also develop a baseline for verifying claims against timeseries evidence and show that even the state-ofthe-art reasoning models like Gemini-2.5-Pro are challenged by time series, achieving a 63.57 accuracy score on verdicts and an Ev 2 R score of 47.36 on verdict justifications.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b04eedb4-00d0-4082-999b-67c8b081b38cCited by top-tier papers3
- ClaimDB: A Fact Verification Benchmark over Large Structured DataMichael Theologitis, Preetam Prabhu Srikar Dammu, Chirag Shah, Dan SuciuACL 2026 · 2 citations
- FakeWorld 1.0: An Omni-modal Benchmark for Fake Media and ContentYifeng Gao, Yifan Ding, Li Wang, Feida Huang et al.ICML 2026
- Beyond Timestamps: Bridging Forward and Backward Reasoning in Temporal Numerical and Relational UnderstandingXinying Qian, Ying Zhang, Xuhui Sui, Yu Zhao et al.ACL 2026
Builds on17
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- The Curious Case of Neural Text DegenerationAri Holtzman, Jan Buys, Li Du, Maxwell Forbes et al.ICLR 2020 · 4,112 citations
- Chatbot Arena: An Open Platform for Evaluating LLMs by Human PreferenceWei-Lin Chiang, Lianmin Zheng, Ying Sheng, Anastasios Nikolas Angelopoulos et al.ICML 2024 · 1,212 citations
- One Fits All: Power General Time Series Analysis by Pretrained LMTian Zhou, Peisong Niu, Xue Wang, Liang Sun et al.NeurIPS 2023 · 1,178 citations
- Time-LLM: Time Series Forecasting by Reprogramming Large Language ModelsMing Jin, Shiyu Wang, Lintao Ma, Zhixuan Chu et al.ICLR 2024 · 915 citations
Related papers
- TimeSeriesExamAgent: Creating Time Series Reasoning Benchmarks at ScaleMalgorzata Gwiazda, Yifu Cai, Mononito Goswami, Arjun Choudhry et al.ICLR 2026 · 6 citations
- TFRBench: A Reasoning Benchmark for Evaluating Forecasting SystemsMd Atik Ahamed, Mihir Parmar, Palash Goyal, Yiwen Song et al.ICML 2026
- HEARTS: Benchmarking LLM Reasoning on Health Time SeriesSirui Li, Shuhan Xiao, Mihir Joshi, Ahmed Metwally et al.ICML 2026 · 7 citations
- FinDVer: Explainable Claim Verification over Long and Hybrid-content Financial DocumentsYilun Zhao, Yitao Long, Tintin Jiang, Chengye Wang et al.EMNLP 2024 · 3 citations
- VeriTaS: The First Dynamic Benchmark for Multimodal Automated Fact-CheckingMark Rothermel, Marcus Kornmann, Marcus Rohrbach, Anna RohrbachACL 2026 · 4 citations
