ContextPRM: Leveraging Contextual Coherence for multi-domain Test-Time Scaling
Haotian Zhang, Liu Liu, Baosheng Yu, Jiayan Qiu, Likang Xiao, Yanwei Ren, Quan Chen, Xianglong Liu
摘要
Process reward models (PRMs) have demonstrated significant efficacy in enhancing the mathematical reasoning capabilities of large language models (LLMs) by leveraging test-time scaling (TTS). However, while most PRMs exhibit substantial gains in mathematical domains, the scarcity of domain-specific training data and knowledge-based learning patterns limits their generalization ability when faced with other domains. To address this limitation, we shift the learning objective from verifying domain-specific knowledge to modeling domainagnostic logical flow. Centering on contextual coherence between chain-of-thought (CoT) steps, our approach is realized through a novel data annotation and training framework, which enhances the model's generalization capabilities across diverse domains. For instance, our resulting model, ContextPRM, achieves a notable 6.5% average accuracy improvement over the majority voting baseline via weighted majority voting across nine nonmathematical domains in MMLU-Pro, including law, history, and philosophy, significantly surpassing the 2.2% improvement from VersaPRM and 0.5% gains from other mathematics-focused PRMs, demonstrating consistent performance across both mathematical and non-mathematical domains. Figure 1. Key results demonstrating ContextPRM's effectiveness. (Left): Training on a single non-math domain surpasses the prior multi-domain SOTA, VersaPRM. (Right): In non-math-adjacent domains, ContextPRM consistently outperforms all baselines under WMV sampling, showcasing superior generalization.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper22
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- Large Language Models are Zero-Shot ReasonersTakeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo 等NeurIPS 2022 · 被引用 8,168 次
- Tree of Thoughts: Deliberate Problem Solving with Large Language ModelsShunyu Yao, Dian Yu, Jeffrey Zhao, Izhak Shafran 等NeurIPS 2023 · 被引用 5,068 次
- Self-Refine: Iterative Refinement with Self-FeedbackAman Madaan, Niket Tandon, Prakhar Gupta, Skyler Hallinan 等NeurIPS 2023 · 被引用 4,972 次
相关 Paper
- VersaPRM: Multi-Domain Process Reward Model via Synthetic Reasoning DataThomas Zeng, Shuibai Zhang, Shutong Wu, Christian Classen 等ICML 2025
- Process Reward Models Meet Planning: Generating Precise and Scalable Datasets for Step-Level RewardsRaffaele Pisano, Roberto NavigliACL 2026 · 被引用 2 次
- Unlocking Multimodal Mathematical Reasoning via Process Reward ModelRuilin Luo, Zhuofan Zheng, Lei Wang, Yifan Wang 等NeurIPS 2025 · 被引用 38 次
- R-PRM: Reasoning-Driven Process Reward ModelingShuaijie She, Junxiao Liu, Yifeng Liu, Jiajun Chen 等EMNLP 2025
- Uncertainty-Based Methods for Automated Process Reward Data Construction and Output Aggregation in Mathematical ReasoningJiuzhou Han, Wray L. Buntine, Ehsan ShareghiAAAI 2026 · 被引用 3 次
