ES4R: Speech Encoding Based on Prepositive Affective Modeling for Empathetic Response Generation
Zhuoyue Gao, Xiaohui Wang, Xiaocui Yang, Wen Zhang, Daling Wang, Shi Feng, Yifei Zhang
Abstract
Empathetic speech dialogue requires not only understanding linguistic content but also perceiving rich paralinguistic information such as prosody, tone, and emotional intensity for affective understandings. Existing speech-to-speech large language models either rely on ASR transcription or use encoders to extract latent representations, often weakening affective information and contextual coherence in multi-turn dialogues. To address this, we propose ES4R, a framework for speech-based empathetic response generation. Our core innovation lies in explicitly modeling structured affective context before speech encoding, rather than relying on implicit learning by the encoder or explicit emotion supervision. Specifically, we introduce a dual-level attention mechanism to capture turn-level affective states and dialogue-level affective dynamics. The resulting affective representations are then integrated with textual semantics through speech-guided cross-modal attention to generate empathetic responses. For speech output, we employ energy-based strategy selection and style fusion to achieve empathetic speech synthesis. ES4R consistently outperforms strong baselines in both automatic and human evaluations and remains robust across different LLM backbones.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 465b13b2-cc8a-4c16-ad5a-3f27e0be79f7Builds on13
- BERTScore: Evaluating Text Generation with BERTTianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger et al.ICLR 2020 · 8,443 citations
- Robust Speech Recognition via Large-Scale Weak SupervisionAlec Radford, Jong Wook Kim, Tao Xu, Greg Brockman et al.ICML 2023 · 6,966 citations
- HuggingGPT: Solving AI Tasks with ChatGPT and its Friends in Hugging FaceYongliang Shen, Kaitao Song, Xu Tan, Dongsheng Li et al.NeurIPS 2023 · 1,778 citations
- SALMONN: Towards Generic Hearing Abilities for Large Language ModelsChangli Tang, Wenyi Yu, Guangzhi Sun, Xianzhao Chen et al.ICLR 2024 · 557 citations
- CEM: Commonsense-Aware Empathetic Response GenerationSahand Sabour, Chujie Zheng, Minlie HuangAAAI 2022 · 196 citations
Related papers
- Talk With Human-like Agents: Empathetic Dialogue Through Perceptible Acoustic Reception and ReactionHaoqiu Yan, Yongxin Zhu, Kai Zheng, Bing Liu et al.ACL 2024
- Knowledge Bridging for Empathetic Dialogue GenerationQintong Li, Piji Li, Zhaochun Ren, Pengjie Ren et al.AAAI 2022 · 128 citations
- Emotionally Situated Text-to-Speech Synthesis in User-Agent ConversationYuchen Liu, Haoyu Zhang, Shichao Liu, Xiang Yin et al.ACM MM 2023 · 6 citations
- STRIDE-ED: A Strategy-Grounded Stepwise Reasoning Framework for Empathetic Dialogue SystemsHongru Ji, Yuyin Fan, Meng Zhao, Xianghua Li et al.ACL 2026 · 1 citation
- UniSS: Unified Expressive Speech-to-Speech Translation with Your VoiceSitong Cheng, Bianweizhen, Xinsheng Wang, Ruibin Yuan et al.ICLR 2026 · 7 citations
