SportsMetrics: Blending Text and Numerical Data to Understand Information Fusion in LLMs
Yebowen Hu, Kaiqiang Song, Sangwoo Cho, Xiaoyang Wang, Hassan Foroosh, Dong Yu, Fei Liu
Abstract
Large language models hold significant potential for integrating various data types, such as text documents and database records, for advanced analytics. However, blending text and numerical data presents substantial challenges. LLMs need to process and cross-reference entities and numbers, handle data inconsistencies and redundancies, and develop planning capabilities such as building a working memory for managing complex data queries. In this paper, we introduce four novel tasks centered around sports data analytics to evaluate the numerical reasoning and information fusion capabilities of LLMs. These tasks involve providing LLMs with detailed, play-by-play sports game descriptions, then challenging them with adversarial scenarios such as new game rules, longer durations, scrambled narratives, and analyzing key statistics in game summaries. We conduct extensive experiments on NBA and NFL games to assess the performance of LLMs on these tasks. Our benchmark, SportsMetrics, introduces a new mechanism for assessing LLMs' numerical reasoning and fusion skills.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 25232e98-bd78-4750-bc95-df75c6541fb1Cited by top-tier papers3
- SPORTU: A Comprehensive Sports Understanding Benchmark for Multimodal Large Language ModelsHaotian Xia, Zhengbang Yang, Junbo Zou, Rhys Tracy et al.ICLR 2025
- Evolving Quantitative Reasoning through Self-Play in Digital Twin MarketsTianmi Ma, Wenxin Huang, Jiawei Du, Lin Li et al.ICML 2026
- Can Large Language Models Derive High-Level Cognition from Low-Level and Fragmented Foundational Information?Yang Liu, Xiaoping Wang, Kai LuAAAI 2025
Builds on16
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou et al.ICLR 2021 · 7,905 citations
- Reflexion: language agents with verbal reinforcement learningNoah Shinn, Federico Cassano, Ashwin Gopinath, Karthik Narasimhan et al.NeurIPS 2023 · 5,828 citations
- Adaptive Chameleon or Stubborn Sloth: Revealing the Behavior of Large Language Models in Knowledge ConflictsJian Xie, Kai Zhang, Jiangjie Chen, Renze Lou et al.ICLR 2024 · 294 citations
- Dense Passage Retrieval for Open-Domain Question AnsweringVladimir Karpukhin, Barlas Oguz, Sewon Min, Patrick Lewis et al.EMNLP 2020 · 142 citations
- Elo Uncovered: Robustness and Best Practices in Language Model EvaluationMeriem Boubdir, Edward Kim, Beyza Ermis, Sara Hooker et al.NeurIPS 2024 · 94 citations
Related papers
- When Reasoning Meets Information Aggregation: A Case Study with Sports NarrativesYebowen Hu, Kaiqiang Song, Sangwoo Cho, Xiaoyang Wang et al.EMNLP 2024 · 10 citations
- RuleArena: A Benchmark for Rule-Guided Reasoning with LLMs in Real-World ScenariosRuiwen Zhou, Wenyue Hua, Liangming Pan, Sitao Cheng et al.ACL 2025 · 13 citations
- BizBench: A Quantitative Reasoning Benchmark for Business and FinanceMichael Krumdick, Rik Koncel-Kedziorski, Viet Dac Lai, Varshini Reddy et al.ACL 2024 · 10 citations
- TaxReasoning: Benchmarking Knowledge-Intensive Mathematical Reasoning with Evolving Tax LawsNan Hu, Yike Wu, Jiaye Li, Huikang Hu et al.AAAI 2026 · 1 citation
- Number Cookbook: Number Understanding of Language Models and How to Improve ItHaotong Yang, Yi Hu, Shijia Kang, Zhouchen Lin et al.ICLR 2025
