UniSRM: A Unified Speech Reward Model for Reasoning-Based Fine-grained Assessment
Yuanyuan Wang, Dongchao Yang, Yayue Deng, Zhiyong Wu, Steven Y. Guo, Helen M. Meng, Xixin Wu
Abstract
Evaluating speech generation still relies heavily on human judgments, such as Mean Opinion Score (MOS), which are expensive, subjective, and difficult to reproduce at scale. While a few recent studies have begun to explore AudioLLM-based judge models, existing efforts typically target only a narrow set of scenarios (e.g., utterance-level quality or singleturn dialogue) and provide limited coverage of diverse speech generation tasks and evaluation dimensions. In this work, we propose UniSRM, a unified speech reward model that can support multi-dimensional, interpretable reward signals with reliable reasoning. To support training and evaluation, we introduce UniSRM-Data and UniSRM-Bench, covering speech evaluation tasks from utterance-level quality to context-level coherence. Based on this dataset, we present the unified speech reward model, UniSRM, with a two-stage pipeline that enables reasoning-based fine-grained assessment. Furthermore, we introduce Reasoning-Consistent Rewards to improve the reliability of the reasoning process. Experiments show that UniSRM delivers more reliable and human-aligned judgments across a broad range of speech evaluation tasks, offering a practical foundation for scalable and unified evaluation of speech quality 1 .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on6
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning et al.NeurIPS 2023 · 10,924 citations
- Unified Multimodal Chain-of-Thought Reward Model through Reinforcement Fine-TuningYibin Wang, Zhimin Li, Yuhang Zang, Chunyu Wang et al.NeurIPS 2025 · 102 citations
- SpeechAlign: Aligning Speech Generation to Human PreferencesDong Zhang, Zhaowei Li, Shimin Li, Xin Zhang et al.NeurIPS 2024 · 74 citations
- DualSpeechLM: Towards Unified Speech Understanding and Generation via Dual Speech Token Modeling with Large Language ModelsYuanyuan Wang, Dongchao Yang, Yiwen Shao, Hangting Chen et al.AAAI 2026 · 3 citations
Related papers
- UniRRM: Unified Reasoning Reward Models Across Languages and Evaluation ParadigmsPeng Lai, Yichao Du, Junchao Wu, Weibo Gao et al.ICML 2026
- SpeechLLM-as-Judges: Towards General and Interpretable Speech Quality EvaluationHui Wang, Jinghua Zhao, Yifan Yang, Shujie Liu et al.ACL 2026 · 21 citations
- SpeechJudge: Towards Human-Level Judgment for Speech NaturalnessXueyao Zhang, Chaoren Wang, Huan Liao, Ziniu Li et al.ICLR 2026 · 32 citations
- Unison: Benchmarking Unified Multimodal Models via Synergistic Understanding and GenerationJinyu Liu, Xincheng Shuai, Henghui Ding, Yu-Gang JiangICML 2026 · 2 citations
- Audio Large Language Models Can Be Descriptive Speech Quality EvaluatorsChen Chen, Yuchen Hu, Siyin Wang, Helin Wang et al.ICLR 2025
