Dual-Axis Generative Reward Model Toward Semantic and Turn-taking Robustness in Interactive Spoken Dialogue Models
Yifu Chen, Shengpeng Ji, Zhengqing Liu, Qian Chen, Wen Wang, Ziqing Wang, Yangzhuo Li, Tianle Liang, Zhou Zhao
Abstract
Achieving seamless, human-like interaction remains a key challenge for full-duplex spoken dialogue models (SDMs). Reinforcement learning (RL) has substantially enhanced text-and vision-language models, while well-designed reward signals are crucial for the performance of RL. We consider RL a promising strategy to address the key challenge for SDMs. However, a fundamental barrier persists: prevailing automated metrics for assessing interaction quality rely on superficial proxies, such as behavioral statistics or timing-prediction accuracy, failing to provide reliable reward signals for RL. On the other hand, human evaluations, despite their richness, remain costly, inconsistent, and difficult to scale. We tackle this critical barrier by proposing a Dual-Axis Generative Reward Model, which is trained to understand complex interaction dynamics using a detailed taxonomy and an annotated dataset, produces a single score and, crucially, provides separate evaluations for semantic quality and interaction timing. Such dual outputs furnish precise diagnostic feedback for SDMs and deliver a dependable, instructive reward signal suitable for online reinforcement learning. Our model achieves state-of-the-art performance on interaction-quality assessment across a wide spectrum of datasets, spanning synthetic dialogues and complex real-world interactions.Our page could be found at https: //github.com/MM-Speech/DualAxisRM .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext bbd31928-c4b3-410c-88cb-66f8e1052a30Builds on10
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning et al.NeurIPS 2023 · 10,924 citations
- SALMONN-omni: A Standalone Speech LLM without Codec Injection for Full-duplex ConversationWenyi Yu, Siyin Wang, Xiaoyu Yang, Xianzhao Chen et al.NeurIPS 2025 · 43 citations
- Align-SLM: Textless Spoken Language Models with Reinforcement Learning from AI FeedbackGuan-Ting Lin, Prashanth Gurunath Shivakumar, Aditya Gourav, Yile Gu et al.ACL 2025 · 33 citations
- Don't Forget Your ABC's: Evaluating the State-of-the-Art in Chat-Oriented Dialogue SystemsSarah E. Finch, James D. Finch, Jinho D. ChoiACL 2023 · 10 citations
- Beyond Turn-Based Interfaces: Synchronous LLMs as Full-Duplex Dialogue AgentsBandhav Veluri, Benjamin N. Peloquin, Bokai Yu, Hongyu Gong et al.EMNLP 2024 · 8 citations
Related papers
- SDiaReward: Modeling and Benchmarking Spoken Dialogue Rewards with Modality and ColloquialnessJingyu Lu, Yuhan Wang, Fan Zhuo, Xize Cheng et al.ACL 2026 · 3 citations
- DUAL RM: Beyond Rule-based Preference Reward Modeling via Meta-RewardXiaobo Liang, Wanfu Wang, Qipeng Huang, Yuyang Ding et al.ACL 2026
- UniSRM: A Unified Speech Reward Model for Reasoning-Based Fine-grained AssessmentYuanyuan Wang, Dongchao Yang, Yayue Deng, Zhiyong Wu et al.ACL 2026
- SageLM: A Multi-aspect and Explainable Large Language Model for Speech JudgementYuan Ge, Junxiang Zhang, Xiaoqian Liu, Bei Li et al.AAAI 2026 · 5 citations
- Aligning Spoken Dialogue Models from User InteractionsAnne Wu, Laurent Mazaré, Neil Zeghidour, Alexandre DéfossezICML 2025
