RuleReasoner: Reinforced Rule-based Reasoning via Domain-aware Dynamic Sampling
Yang Liu, Jiaqi Li, Zilong Zheng
Abstract
Rule-based reasoning is acknowledged as one of the fundamental problems of reasoning. While recent studies show that large reasoning models (LRMs) have remarkable reasoning capabilities enhanced by reinforcement learning (RL), real applications still face severe challenges due to variations in rule formats, types, and complexity. To mitigate this issue, we introduce RULEREASONER, an effective method for rule-based reasoning via a wide collection of curated tasks and a novel domain-aware dynamic sampling approach in RL. Specifically, RULEREASONER resamples each training batch by updating the domain weights based on historical rewards. This facilitates domain balance and active learning schedules for RL, obviating static mix-training engineered by human. Evaluations on in-distribution (ID) and out-of-distribution (OOD) benchmarks reveal that RULEREASONER outperforms frontier LRMs by a significant margin (∆4.1% on eight ID tasks and ∆10.4% on three OOD tasks over OpenAI-o1). Notably, our approach also exhibits higher computational efficiency compared to prior methods. Recent work has demonstrated the remarkable reasoning capabilities of large reasoning models (LRMs) with an intermediate thinking process, chain-of-thought (CoT) (Wei et al., 2022b), notably ⋆ Equal contribution. † Corresponding author.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 6bc2981c-02e2-4e77-a9a8-4228adee4d32Cited by top-tier papers4
- Native Parallel Reasoner: Reasoning in Parallelism via Self-Distilled Reinforcement LearningTong Wu, Michael Liu, Jun Bai, Zixia Jia et al.ICML 2026 · 12 citations
- LogicReward: Incentivizing LLM Reasoning via Step-Wise Logical SupervisionJundong Xu, Hao (Scofield) Fei, Huichi Zhou, Xin Quan et al.ICLR 2026 · 9 citations
- Reinforced Query Reasoners for Reasoning-intensive Retrieval TasksXubo Qin, Jun Bai, Jiaqi Li, Zixia Jia et al.EMNLP 2025
- Understanding and Leveraging the Expert Specialization of Context Faithfulness in Mixture-of-Experts LLMsJun Bai, Minghao Tong, Yang Liu, Zixia Jia et al.EMNLP 2025
Builds on31
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- Finetuned Language Models are Zero-Shot LearnersJason Wei, Maarten Bosma, Vincent Y. Zhao, Kelvin Guu et al.ICLR 2022 · 4,966 citations
- DAPO: An Open-Source LLM Reinforcement Learning System at ScaleQiying Yu, Zheng Zhang, Ruofei Zhu, Yufeng Yuan et al.NeurIPS 2025 · 2,828 citations
- Efficient Memory Management for Large Language Model Serving with PagedAttentionWoosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng et al.SOSP 2023 · 1,016 citations
- Self-Consistency Improves Chain of Thought Reasoning in Language ModelsXuezhi Wang, Jason Wei, Dale Schuurmans, Quoc V. Le et al.ICLR 2023 · 681 citations
Related papers
- Training Language Models to Reason EfficientlyDaman Arora, Andrea ZanetteNeurIPS 2025 · 270 citations
- General-Reasoner: Advancing LLM Reasoning Across All DomainsXueguang Ma, Qian Liu, Dongfu Jiang, Ge Zhang et al.NeurIPS 2025 · 153 citations
- How Far Are We from Optimal Reasoning Efficiency?Jiaxuan Gao, Shu Yan, Qixin Tan, Lu Yang et al.NeurIPS 2025 · 12 citations
- AdaReason: Progressive Training of Multi-LoRA Adapters for Budget-Adaptive Language Reasoning ModelsJiacheng Wang, Tianle Chen, Pengyu Cheng, Xiaofeng Hou et al.AAAI 2026
- Does Your Reasoning Model Implicitly Know When to Stop Thinking?Zixuan Huang, Xin Xia, Yuxi Ren, Jianbin Zheng et al.ICML 2026 · 21 citations
