SECap: Speech Emotion Captioning with Large Language Model
Yaoxun Xu, Hangting Chen, Jianwei Yu, Qiaochu Huang, Zhiyong Wu, Shi-Xiong Zhang, Guangzhi Li, Yi Luo, Rongzhi Gu
摘要
Speech emotions are crucial in human communication and are extensively used in fields like speech synthesis and natural language understanding. Most prior studies, such as speech emotion recognition, have categorized speech emotions into a fixed set of classes. Yet, emotions expressed in human speech are often complex, and categorizing them into predefined groups can be insufficient to adequately represent speech emotions. On the contrary, describing speech emotions directly by means of natural language may be a more effective approach. Regrettably, there are not many studies available that have focused on this direction. Therefore, this paper proposes a speech emotion captioning framework named SECap, aiming at effectively describing speech emotions using natural language. Owing to the impressive capabilities of large language models in language comprehension and text generation, SECap employs LLaMA as the text decoder to allow the production of coherent speech emotion captions. In addition, SECap leverages HuBERT as the audio encoder to extract general speech features and Q-Former as the Bridge-Net to provide LLaMA with emotion-related speech features. To accomplish this, Q-Former utilizes mutual information learning to disentangle emotion-related speech features and speech contents, while implementing contrastive learning to extract more emotion-related speech features. The results of objective and subjective evaluations demonstrate that: 1) the SE-Cap framework outperforms the HTSAT-BART baseline in all objective evaluations; 2) SECap can generate high-quality speech emotion captions that attain performance on par with human annotators in subjective mean opinion score tests.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper12
- Language Model Can Listen While SpeakingZiyang Ma, Yakun Song, Chenpeng Du, Jian Cong 等AAAI 2025 · 被引用 58 次
- EmotionThinker: Prosody-Aware Reinforcement Learning for Explainable Speech Emotion ReasoningDingdong WANG, Shujie LIU, Tianhua Zhang, Youjun Chen 等ICLR 2026 · 被引用 22 次
- QualiSpeech: A Speech Quality Assessment Dataset with Natural Language Reasoning and DescriptionsSiyin Wang, Wenyi Yu, Xianzhao Chen, Xiaohai Tian 等ACL 2025 · 被引用 20 次
- EmoVoice: LLM-based Emotional Text-To-Speech Model with Freestyle Text PromptingGuanrou Yang, Chen Yang, Qian Chen, Ziyang Ma 等ACM MM 2025 · 被引用 16 次
- SpeechCraft: A Fine-Grained Expressive Speech Dataset with Natural Language DescriptionZeyu Jin, Jia Jia, Qixin Wang, Kehan Li 等ACM MM 2024 · 被引用 12 次
它引用的顶会 Paper5
- Large Language Models are Zero-Shot ReasonersTakeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo 等NeurIPS 2022 · 被引用 8,168 次
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 被引用 7,873 次
- HuggingGPT: Solving AI Tasks with ChatGPT and its Friends in Hugging FaceYongliang Shen, Kaitao Song, Xu Tan, Dongsheng Li 等NeurIPS 2023 · 被引用 1,778 次
- CLUB: A Contrastive Log-ratio Upper Bound of Mutual InformationPengyu Cheng, Weituo Hao, Shuyang Dai, Jiachang Liu 等ICML 2020 · 被引用 512 次
- GLM: General Language Model Pretraining with Autoregressive Blank InfillingZhengxiao Du, Yujie Qian, Xiao Liu, Ming Ding 等ACL 2022
相关 Paper
- AlignCap: Aligning Speech Emotion Captioning to Human PreferencesZiqi Liang, Haoxiang Shi, Hanhui ChenEMNLP 2024 · 被引用 2 次
- Voices, Faces, and Feelings: Multi-modal Emotion-Cognition Captioning for Mental Health UnderstandingZhiyuan Zhou, Yanrong Guo, Shijie HaoAAAI 2026
- video-SALMONN: Speech-Enhanced Audio-Visual Large Language ModelsGuangzhi Sun, Wenyi Yu, Changli Tang, Xianzhao Chen 等ICML 2024 · 被引用 92 次
- AutoAD III: The Prequel - Back to the PixelsTengda Han, Max Bain, Arsha Nagrani, Gül Varol 等CVPR 2024
- DQ-Former: Querying Transformer with Dynamic Modality Priority for Cognitive-aligned Multimodal Emotion Recognition in ConversationJing Ye, Xinpei ZhaoACM MM 2024 · 被引用 10 次
