Reasoning over Uncertain Text by Generative Large Language Models
Aliakbar Nafar, Kristen Brent Venable, Parisa Kordjamshidi
摘要
This paper considers the challenges Large Language Models (LLMs) face when reasoning over text that includes information involving uncertainty explicitly quantified via probability values. This type of reasoning is relevant to a variety of contexts ranging from everyday conversations to medical decision-making. Despite improvements in the mathematical reasoning capabilities of LLMs, they still exhibit significant difficulties when it comes to probabilistic reasoning. To deal with this problem, we introduce the Bayesian Linguistic Inference Dataset (BLInD), a new dataset specifically designed to test the probabilistic reasoning capabilities of LLMs. We use BLInD to find out the limitations of LLMs for tasks involving probabilistic reasoning. In addition, we present several prompting strategies that map the problem to different formal representations, including Python code, probabilistic algorithms, and probabilistic logical programming. We conclude by providing an evaluation of our methods on BLInD and an adaptation of a causal reasoning question-answering dataset. Our empirical results highlight the effectiveness of our proposed strategies for multiple LLMs.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper9
- Toolformer: Language Models Can Teach Themselves to Use ToolsTimo Schick, Jane Dwivedi-Yu, Roberto Dessì, Roberta Raileanu 等NeurIPS 2023 · 被引用 5,989 次
- NumGLUE: A Suite of Fundamental yet Challenging Mathematical Reasoning TasksSwaroop Mishra, Arindam Mitra, Neeraj Varshney, Bhavdeep Singh Sachdeva 等ACL 2022 · 被引用 138 次
- StepGame: A New Benchmark for Robust Multi-Hop Spatial Reasoning in TextsZhengxiang Shi, Qiang Zhang, Aldo LipaniAAAI 2022 · 被引用 100 次
- CLadder: A Benchmark to Assess Causal Reasoning Capabilities of Language ModelsZhijing Jin, Yuen Chen, Felix Leeb, Luigi Gresele 等NeurIPS 2023 · 被引用 74 次
- GLUECons: A Generic Benchmark for Learning under ConstraintsHossein Rajaby Faghihi, Aliakbar Nafar, Chen Zheng, Roshanak Mirzaee 等AAAI 2023 · 被引用 18 次
相关 Paper
- Reasoning While Asking: Transforming Reasoning Large Language Models from Passive Solvers to Proactive InquirersXin Chen, Feng Jiang, Yiqian Zhang, Hardy Chen 等ACL 2026
- QUITE: Quantifying Uncertainty in Natural Language Text in Bayesian Reasoning ScenariosTimo Pierre Schrader, Lukas Lange, Simon Razniewski, Annemarie FriedrichEMNLP 2024
- Can Large Language Models Infer Causation from Correlation?Zhijing Jin, Jiarui Liu, Zhiheng Lyu, Spencer Poff 等ICLR 2024 · 被引用 186 次
- SciBench: Evaluating College-Level Scientific Problem-Solving Abilities of Large Language ModelsXiaoxuan Wang, Ziniu Hu, Pan Lu, Yanqiao Zhu 等ICML 2024 · 被引用 220 次
- NoisyCausal: A Benchmark for Evaluating Causal Reasoning Under Structured NoiseZhi Xu, Yun FuACL 2026
