Beyond Timestamps: Bridging Forward and Backward Reasoning in Temporal Numerical and Relational Understanding
Xinying Qian, Ying Zhang, Xuhui Sui, Yu Zhao, Baohang Zhou, Jeff Z. Pan
摘要
Temporal reasoning remains a critical challenge for large language models (LLMs), particularly when it requires encompassing relational dependencies and numerical constraints. Yet, existing benchmarks largely overlook the joint consideration of these two dimensions and primarily rely on single-task evaluation paradigms, making it difficult to assess whether correct answers reflect grounded reasoning or arise from superficial statistical recall. To address these gaps, we introduce TNR, a benchmark designed to evaluate both Temporal Numerical and Relational reasoning. We propose a bi-directional evaluation framework consisting of forward generation via Question Answering (QA) and backward verification via Fact Verification (FV). By measuring the alignment between QA and FV, we introduce a Consistency Rate to quantify the robustness of reasoning across these two directions. Experiments on a range of LLMs reveal notable discrepancies between QA and FV performance, particularly in numerical and interval-based tasks. Moreover, our bi-directional error analysis demonstrates that these inconsistencies often stem from heuristic shortcuts and statistical co-occurrences rather than grounded logical deduction, flaws that are frequently masked in standard single-task evaluations.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper15
- TabFact: A Large-scale Dataset for Table-based Fact VerificationWenhu Chen, Hongmin Wang, Jianshu Chen, Yunkai Zhang 等ICLR 2020 · 被引用 674 次
- SelfCheckGPT: Zero-Resource Black-Box Hallucination Detection for Generative Large Language ModelsPotsawee Manakul, Adian Liusie, Mark J. F. GalesEMNLP 2023 · 被引用 331 次
- MathFusion: Enhancing Mathematical Problem-solving of LLM through Instruction FusionQizhi Pei, Lijun Wu, Zhuoshi Pan, Yu Li 等ACL 2025 · 被引用 27 次
- Towards Benchmarking and Improving the Temporal Reasoning Capability of Large Language ModelsQingyu Tan, Hwee Tou Ng, Lidong BingACL 2023 · 被引用 24 次
- UnSeenTimeQA: Time-Sensitive Question-Answering Beyond LLMs' MemorizationMd Nayem Uddin, Amir Saeidi, Divij Handa, Agastya Seth 等ACL 2025 · 被引用 18 次
相关 Paper
- Harnessing Temporal Databases for Systematic Evaluation of Factual Time-Sensitive Question-Answering in LLMsSoyeon Kim, Jindong Wang, Xing Xie, Steven Euijong WhangICLR 2026
- Do Language Models Have a Common Sense regarding Time? Revisiting Temporal Commonsense Reasoning in the Era of Large Language ModelsRaghav Jain, Daivik Sojitra, Arkadeep Acharya, Sriparna Saha 等EMNLP 2023 · 被引用 17 次
- Temporal Referential Consistency: Do LLMs Favor Sequences Over Absolute Time References?Ashutosh Bajpai, Tanmoy ChakrabortyEMNLP 2025
- Reasoning Gets Harder for LLMs Inside A DialogueIvan Kartác, Mateusz Lango, Ondrej DusekACL 2026 · 被引用 2 次
- Living in the Moment: Can Large Language Models Grasp Co-Temporal Reasoning?Zhaochen Su, Juntao Li, Jun Zhang, Tong Zhu 等ACL 2024
