RuleArena: A Benchmark for Rule-Guided Reasoning with LLMs in Real-World Scenarios
Ruiwen Zhou, Wenyue Hua, Liangming Pan, Sitao Cheng, Xiaobao Wu, En Yu, William Yang Wang
Abstract
This paper introduces RuleArena, a novel and challenging benchmark designed to evaluate the ability of large language models (LLMs) to follow complex, real-world rules in reasoning. Covering three practical domains -- airline baggage fees, NBA transactions, and tax regulations -- RuleArena assesses LLMs' proficiency in handling intricate natural language instructions that demand long-context understanding, logical reasoning, and accurate mathematical computation. Two key attributes distinguish RuleArena from traditional rule-based reasoning benchmarks: (1) it extends beyond standard first-order logic representations, and (2) it is grounded in authentic, practical scenarios, providing insights into the suitability and reliability of LLMs for real-world applications. Our findings reveal several notable limitations in LLMs: (1) they struggle to identify and apply the appropriate rules, frequently becoming confused by similar but distinct regulations, (2) they cannot consistently perform accurate mathematical computations, even when they correctly identify the relevant rules, and (3) in general, they perform poorly in the benchmark. We also observe a significant performance boost when LLMs are provided with external tools for oracle math and logic operations. These results highlight significant challenges and promising research directions in advancing LLMs' rule-guided reasoning capabilities in real-life applications. Our codes and data are publicly available on https://github.com/skyriver-2000/RuleArena.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 8f85a3d1-c828-4a29-8b9a-4a22961d3c09Cited by top-tier papers6
- AntiLeakBench: Preventing Data Contamination by Automatically Constructing Benchmarks with Updated Real-World KnowledgeXiaobao Wu, Liangming Pan, Yuxi Xie, Ruiwen Zhou et al.ACL 2025 · 35 citations
- Beyond Single-Task: Robust Multi-Task Length Generalization for LLMsYi Hu, Shijia Kang, Haotong Yang, Haotian Xu et al.NeurIPS 2025 · 6 citations
- Making Logic a First-Class Citizen in Generative ML for NetworkingHongyu Hè, Minhao Jin, Maria ApostolakiNSDI 2026 · 5 citations
- MuSLR: Multimodal Symbolic Logical ReasoningJundong Xu, Hao Fei, Yuhui Zhang, Liangming Pan et al.NeurIPS 2025 · 5 citations
- Towards Advanced Mathematical Reasoning for LLMs via First-Order Logic Theorem ProvingChuxue Cao, Mengze Li, Juntao Dai, Jinluan Yang et al.EMNLP 2025 · 1 citation
Builds on12
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- Large Language Models are Zero-Shot ReasonersTakeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo et al.NeurIPS 2022 · 8,168 citations
- Interleaving Retrieval with Chain-of-Thought Reasoning for Knowledge-Intensive Multi-Step QuestionsHarsh Trivedi, Niranjan Balasubramanian, Tushar Khot, Ashish SabharwalACL 2023 · 187 citations
- AntiLeakBench: Preventing Data Contamination by Automatically Constructing Benchmarks with Updated Real-World KnowledgeXiaobao Wu, Liangming Pan, Yuxi Xie, Ruiwen Zhou et al.ACL 2025 · 35 citations
- IDEAL: Influence-Driven Selective Annotations Empower In-Context Learners in Large Language ModelsShaokun Zhang, Xiaobo Xia, Zhaoqing Wang, Ling-Hao Chen et al.ICLR 2024 · 28 citations
Related papers
- TaxReasoning: Benchmarking Knowledge-Intensive Mathematical Reasoning with Evolving Tax LawsNan Hu, Yike Wu, Jiaye Li, Huikang Hu et al.AAAI 2026 · 1 citation
- GraphArena: Evaluating and Exploring Large Language Models on Graph ComputationJianheng Tang, Qifan Zhang, Yuhan Li, Nuo Chen et al.ICLR 2025
- GuessArena: Guess Who I Am? A Self-Adaptive Framework for Evaluating LLMs in Domain-Specific Knowledge and ReasoningQingchen Yu, Zifan Zheng, Ding Chen, Simin Niu et al.ACL 2025 · 5 citations
- GameArena: Evaluating LLM Reasoning through Live Computer GamesLanxiang Hu, Qiyu Li, Anze Xie, Nan Jiang et al.ICLR 2025
- SC-Arena: A Natural Language Benchmark for Single-Cell Reasoning with Knowledge-Augmented EvaluationJiahao Zhao, Feng Jiang, Shaowei Qin, Zhonghui Zhang et al.ICLR 2026 · 4 citations
