CLARity: Reasoning Consistency Alone Can Teach Reinforced Experts
Jiuheng Lin, Cong Jiang, Zirui Wu, Jiarui Sun, Yansong Feng
摘要
Training expert LLMs in domains with scarce fine-grained annotated data is admittedly challenging, often relying on multiple-choice questions (MCQs). However, standard outcomebased reinforcement learning (RL) on MCQs is risky. While outcome-based RL may improve accuracy, it frequently compromises the reasoning process, yielding internally inconsistent rationales that diverge from the final predictions. Existing solutions to supervise the reasoning process, such as large-scale Process Reward Models (PRMs), are prohibitively expensive. To address this, we propose CLARITY, a costeffective RL framework that enhances reasoning quality using a small, general-purpose LLM only. CLARITY integrates a consistency-aware reward mechanism with a 2-stage refine-thenmonitor training pipeline to enhance reasoning consistency, and a dynamic data reformulation strategy to better exploit annotated data available. Experiments demonstrate that CLARITY can improve the consistency of responses by 16.5% over standard outcome-based RL, and bring an improvement of 7.5% in final accuracy. Human evaluations further confirm substantial gains in factual correctness and reasoning coherence, leading to more trustworthy model outputs. Thus, CLARITY offers a generalizable solution that enables smaller models to effectively guide expert LLM training by monitoring reasoning consistency. 1
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper6
- Efficient Memory Management for Large Language Model Serving with PagedAttentionWoosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng 等SOSP 2023 · 被引用 1,016 次
- ReTool: Reinforcement Learning for Strategic Tool Use in LLMsJiazhan Feng, Shijue Huang, Xingwei Qu, Ge Zhang 等ICLR 2026 · 被引用 406 次
- JEC-QA: A Legal-Domain Question Answering DatasetHaoxi Zhong, Chaojun Xiao, Cunchao Tu, Tianyang Zhang 等AAAI 2020 · 被引用 212 次
- Chain-of-Thought Reasoning In The Wild Is Not Always FaithfulIván Arcuschin, Jett Janiak, Robert Krzyzanowski, Senthooran Rajamanoharan 等ICML 2026 · 被引用 175 次
- HybridFlow: A Flexible and Efficient RLHF FrameworkGuangming Sheng, Chi Zhang, Zilingfeng Ye, Xibin Wu 等EuroSys 2025 · 被引用 61 次
相关 Paper
- KnowRL: Exploring Knowledgeable Reinforcement Learning for FactualityBaochang Ren, Shuofei Qiao, Ningyu Zhang, Da Zheng 等ACL 2026 · 被引用 12 次
- Distributional Clarity: The Hidden Driver of RL-Friendliness in Large Language ModelsShaoning Sun, Mingzhu Cai, Huang He, Bingjin Chen 等ACL 2026 · 被引用 1 次
- Enhancing the Outcome Reward-based RL Training of MLLMs with Self-Consistency SamplingJiahao Wang, Weiye Xu, Aijun Yang, Wengang Zhou 等NeurIPS 2025 · 被引用 4 次
- Monitorability as a Free Gift: How RLVR Spontaneously Aligns ReasoningZidi Xiong, Shan Chen, Himabindu LakkarajuICML 2026 · 被引用 3 次
- Act as you think: Reinforcing Consistent Reasoning in Medical Visual Question AnsweringSongtao Jiang, Yuan Wang, Ruizhe Chen, Yan Zhang 等ACL 2026
