When Reasoning Meets Information Aggregation: A Case Study with Sports Narratives
Yebowen Hu, Kaiqiang Song, Sangwoo Cho, Xiaoyang Wang, Wenlin Yao, Hassan Foroosh, Dong Yu, Fei Liu
摘要
Reasoning is most powerful when an LLM accurately aggregates relevant information. We examine the critical role of information aggregation in reasoning by requiring the LLM to analyze sports narratives. To succeed at this task, an LLM must infer points from actions, identify related entities, attribute points accurately to players and teams, and compile key statistics to draw conclusions. We conduct comprehensive experiments with real NBA basketball data and present SPORTSGEN, a new method to synthesize game narratives. By synthesizing data, we can rigorously evaluate LLMs' reasoning capabilities under complex scenarios with varying narrative lengths and density of information. Our findings show that most models, including GPT-4o, often fail to accurately aggregate basketball scores due to frequent scoring patterns. Open-source models like Llama-3 further suffer from significant score hallucinations. Finally, the effectiveness of reasoning is influenced by narrative complexity, information density, and domain-specific terms, highlighting the challenges in analytical reasoning tasks. 1 * Work done during Yebowen Hu's internship; Kaiqiang Song and Sangwoo Cho were full-time researchers at Tencent AI Lab, Seattle, USA at the time of this work.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Benchmarking Large Language Models Under Data Contamination: A Survey from Static to Dynamic EvaluationSimin Chen, Yiming Chen, Zexin Li, Yifan Jiang 等EMNLP 2025 · 被引用 2 次
- SPORTU: A Comprehensive Sports Understanding Benchmark for Multimodal Large Language ModelsHaotian Xia, Zhengbang Yang, Junbo Zou, Rhys Tracy 等ICLR 2025
- MIRAGE: Evaluating and Explaining Inductive Reasoning Process in Language ModelsJiachun Li, Pengfei Cao, Zhuoran Jin, Yubo Chen 等ICLR 2025
它引用的顶会 Paper16
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- The Reversal Curse: LLMs trained on "A is B" fail to learn "B is A"Lukas Berglund, Meg Tong, Maximilian Kaufmann, Mikita Balesni 等ICLR 2024 · 被引用 462 次
- Chain-of-Thought Reasoning Without PromptingXuezhi Wang, Denny ZhouNeurIPS 2024 · 被引用 305 次
- Take a Step Back: Evoking Reasoning via Abstraction in Large Language ModelsHuaixiu Steven Zheng, Swaroop Mishra, Xinyun Chen, Heng-Tze Cheng 等ICLR 2024 · 被引用 216 次
- MuSR: Testing the Limits of Chain-of-thought with Multistep Soft ReasoningZayne Sprague, Xi Ye, Kaj Bostrom, Swarat Chaudhuri 等ICLR 2024 · 被引用 172 次
相关 Paper
- SportsMetrics: Blending Text and Numerical Data to Understand Information Fusion in LLMsYebowen Hu, Kaiqiang Song, Sangwoo Cho, Xiaoyang Wang 等ACL 2024 · 被引用 6 次
- STARQA: A Question Answering Dataset for Complex Analytical Reasoning over Structured DatabasesMounica Maddela, Lingjue Xie, Daniel Preotiuc-Pietro, MausamEMNLP 2025
- One Thousand and One Pairs: A "novel" challenge for long-context language modelsMarzena Karpinska, Katherine Thai, Kyle Lo, Tanya Goyal 等EMNLP 2024 · 被引用 6 次
- Inferring Events from Time Series using Language ModelsMingtian Tan, Mike A. Merrill, Zachary Gottesman, Tim Althoff 等ACL 2026 · 被引用 7 次
- Sportify: Question Answering with Embedded Visualizations and Personified Narratives for Sports VideoChunggi Lee, Tica Lin, Hanspeter Pfister, Chen Zhu-TianIEEE VIS 2024 · 被引用 7 次
