Reward Modeling for Scientific Writing Evaluation
Furkan Sahinuç, Subhabrata Dutta, Iryna Gurevych
Abstract
Scientific writing is an expert-domain task that demands deep domain knowledge, task-specific requirements and reasoning capabilities that leverage the domain knowledge to satisfy the task specifications. While scientific text generation has been widely studied, its evaluation remains a challenging and open problem. It is critical to develop models that can be reliably deployed for evaluating diverse openended scientific writing tasks while adhering to their distinct requirements. However, existing LLM-based judges and reward models are primarily optimized for general-purpose benchmarks with fixed scoring rubrics and evaluation criteria. Consequently, they often fail to reason over sparse knowledge of scientific domains when interpreting task-dependent and multi-faceted criteria. Moreover, fine-tuning for each individual task is costly and impractical for low-resource settings. To bridge these gaps, we propose cost-efficient, open-source reward models tailored for scientific writing evaluation. We introduce a two-stage training framework that initially optimizes scientific evaluation preferences and then refines reasoning capabilities. Our multi-aspect evaluation design and joint training across diverse tasks enable fine-grained assessment and robustness to dynamic criteria and scoring rubrics. Experimental analysis shows that our training regime strongly improves LLM-based scientific writing evaluation. Our models generalize effectively across tasks and to previously unseen scientific writing evaluation settings, allowing a single trained evaluator to be reused without task-specific retraining. We make our code 1 and data 2 publicly available.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext f06030f3-0c6d-464f-a93d-95fd059d505aBuilds on10
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Defining and Characterizing Reward GamingJoar Skalse, Nikolaus H. R. Howe, Dmitrii Krasheninnikov, David KruegerNeurIPS 2022 · 466 citations
- Reinforcement Learning with Verifiable Rewards Implicitly Incentivizes Correct Reasoning in Base LLMsXumeng Wen, Zihan Liu, Shun Zheng, Shengyu Ye et al.ICLR 2026 · 279 citations
- RM-R1: Reward Modeling as ReasoningXiusi Chen, Gaotang Li, Ziqi Wang, Bowen Jin et al.ICLR 2026 · 147 citations
- J1: Incentivizing Thinking in LLM-as-a-Judge via Reinforcement LearningChenxi Whitehouse, Tianlu Wang, Ping Yu, Xian Li et al.ICLR 2026 · 74 citations
Related papers
- UniRRM: Unified Reasoning Reward Models Across Languages and Evaluation ParadigmsPeng Lai, Yichao Du, Junchao Wu, Weibo Gao et al.ICML 2026
- Triviality Corrected Endogenous RewardXinda Wang, Zhengxu Hou, Yangshijie Zhang, Bingren Yan et al.ACL 2026
- SciEx: Benchmarking Large Language Models on Scientific Exams with Human Expert Grading and Automatic GradingTu Anh Dinh, Carlos Mullov, Leonard Bärmann, Zhaolin Li et al.EMNLP 2024 · 10 citations
- mR3: Multilingual Rubric-Agnostic Reward Reasoning ModelsDavid Anugraha, Shou-Yi Hung, Zilu Tang, En-Shiun Annie Lee et al.ICLR 2026 · 9 citations
- Sci2Pol: Evaluating and Fine-tuning LLMs on Scientific-to-Policy Brief GenerationWeimin Wu, Alexander C. Furnas, Eddie Yang, Gefei Liu et al.ICLR 2026 · 1 citation
