Inferring Emotion from Large-scale Internet Voice Data: A Semi-supervised Curriculum Augmentation based Deep Learning Approach
Suping Zhou, Jia Jia, Zhiyong Wu, Zhihan Yang, Yanfeng Wang, Wei Chen, Fanbo Meng, Shuo Huang, Jialie Shen, Xiaochuan Wang
Abstract
Effective emotion inference from user queries helps to give a more personified response for Voice Dialogue Applications(VDAs). The tremendous amounts of VDA users bring in diverse emotion expressions. How to achieve a high emotion inferring performance from large-scale Internet Voice Data in VDAs? Traditionally, researches on speech emotion recognition are based on acted voice datasets, which have limited speakers but strong and clear emotion expressions. Inspired by this, in this paper, we propose a novel approach to leverage acted voice data with strong emotion expressions to enhance large-scale unlabeled internet voice data with diverse emotion expressions for emotion inferring. Specifically, we propose a novel semi-supervised multi-modal curriculum augmentation deep learning framework. First, to learn more general emotion cues, we adopt a curriculum learning based epoch-wise training strategy, which trains our model guided by strong and balanced emotion samples from acted voice data and sub-sequently leverages weak and unbalanced emotion samples from internet voice data.Second, to employ more diverse emotion expressions, we design a Multi-path Mix-match Multimodal Deep Neural Network(MMMD), which effectively learns feature representations for multiple modalities and trains labeled and unlabeled data in hybrid semi-supervised methods for superior generalization and robustness. Experiments on an internet voice dataset with 500,000 utterances show our method outperforms (+10.09% in terms of F1) several alternative baselines, while an acted corpus with 2,397 utterances contributes 4.35%. To further compare our method with state-of-the-art techniques in traditionally acted voice datasets, we also conduct experiments on public dataset IEMOCAP. The results reveal the effectiveness of the proposed approach.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext eef1a4a9-d8a3-4019-8d31-e3ef614c9b77Cited by top-tier papers2
- Robust Multimodal Sentiment Analysis of Image-Text Pairs by Distribution-Based Feature Recovery and FusionDaiqing Wu, Dongbao Yang, Yu Zhou, Can MaACM MM 2024 · 13 citations
- ASE: Practical Acoustic Speed Estimation Beyond Doppler via Sound Diffusion FieldSheng Lyu, Chenshu WuUbiComp 2025 · 4 citations
Related papers
- Semi-supervised Multi-modal Emotion Recognition with Cross-Modal Distribution MatchingJingjun Liang, Ruichen Li, Qin JinACM MM 2020 · 67 citations
- Towards Emotion-aided Multi-modal Dialogue Act ClassificationTulika Saha, Aditya Prakash Patra, Sriparna Saha, Pushpak BhattacharyyaACL 2020 · 63 citations
- AU-assisted Graph Attention Convolutional Network for Micro-Expression RecognitionHong-Xia Xie, Ling Lo, Hong-Han Shuai, Wen-Huang ChengACM MM 2020 · 189 citations
- MMGCN: Multimodal Fusion via Deep Graph Convolution Network for Emotion Recognition in ConversationJingwen Hu, Yuchen Liu, Jinming Zhao, Qin JinACL 2021
- Do LLMs Feel? Teaching Emotion Recognition with Prompts, Retrieval, and Curriculum LearningXinran Li, Yu Liu, Jiaqi Qiao, Xiujuan XuAAAI 2026 · 1 citation
