TSVer: A Benchmark for Fact Verification Against Time-Series Evidence
Marek Strong, Andreas Vlachos
摘要
Reasoning over temporal and numerical data, such as time series, is a crucial aspect of factchecking. While many systems have recently been developed to handle this form of evidence, their evaluation remains limited by existing datasets, which often lack structured evidence, provide insufficient justifications for verdicts, or rely on synthetic claims. In this paper, we introduce TSVER, a new benchmark dataset for fact verification focusing on temporal and numerical reasoning with time-series evidence. TSVER contains 304 real-world claims sourced from 41 fact-checking organizations and a curated database of 400 time series covering diverse domains. Each claim is annotated with time frames across all pertinent time series, along with a verdict and justifications reflecting how the evidence is used to reach the verdict. Using an LLM-assisted multi-step annotation process, we improve the quality of our annotations and achieve an inter-annotator agreement of κ = 0.77 on verdicts. We also develop a baseline for verifying claims against timeseries evidence and show that even the state-ofthe-art reasoning models like Gemini-2.5-Pro are challenged by time series, achieving a 63.57 accuracy score on verdicts and an Ev 2 R score of 47.36 on verdict justifications.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- ClaimDB: A Fact Verification Benchmark over Large Structured DataMichael Theologitis, Preetam Prabhu Srikar Dammu, Chirag Shah, Dan SuciuACL 2026 · 被引用 2 次
- FakeWorld 1.0: An Omni-modal Benchmark for Fake Media and ContentYifeng Gao, Yifan Ding, Li Wang, Feida Huang 等ICML 2026
- Beyond Timestamps: Bridging Forward and Backward Reasoning in Temporal Numerical and Relational UnderstandingXinying Qian, Ying Zhang, Xuhui Sui, Yu Zhao 等ACL 2026
它引用的顶会 Paper17
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- The Curious Case of Neural Text DegenerationAri Holtzman, Jan Buys, Li Du, Maxwell Forbes 等ICLR 2020 · 被引用 4,112 次
- Chatbot Arena: An Open Platform for Evaluating LLMs by Human PreferenceWei-Lin Chiang, Lianmin Zheng, Ying Sheng, Anastasios Nikolas Angelopoulos 等ICML 2024 · 被引用 1,212 次
- One Fits All: Power General Time Series Analysis by Pretrained LMTian Zhou, Peisong Niu, Xue Wang, Liang Sun 等NeurIPS 2023 · 被引用 1,178 次
- Time-LLM: Time Series Forecasting by Reprogramming Large Language ModelsMing Jin, Shiyu Wang, Lintao Ma, Zhixuan Chu 等ICLR 2024 · 被引用 915 次
相关 Paper
- TimeSeriesExamAgent: Creating Time Series Reasoning Benchmarks at ScaleMalgorzata Gwiazda, Yifu Cai, Mononito Goswami, Arjun Choudhry 等ICLR 2026 · 被引用 6 次
- TFRBench: A Reasoning Benchmark for Evaluating Forecasting SystemsMd Atik Ahamed, Mihir Parmar, Palash Goyal, Yiwen Song 等ICML 2026
- HEARTS: Benchmarking LLM Reasoning on Health Time SeriesSirui Li, Shuhan Xiao, Mihir Joshi, Ahmed Metwally 等ICML 2026 · 被引用 7 次
- FinDVer: Explainable Claim Verification over Long and Hybrid-content Financial DocumentsYilun Zhao, Yitao Long, Tintin Jiang, Chengye Wang 等EMNLP 2024 · 被引用 3 次
- VeriTaS: The First Dynamic Benchmark for Multimodal Automated Fact-CheckingMark Rothermel, Marcus Kornmann, Marcus Rohrbach, Anna RohrbachACL 2026 · 被引用 4 次
