The MERSA Dataset and a Transformer-Based Approach for Speech Emotion Recognition
Enshi Zhang, Rafael Trujillo, Christian Poellabauer
Abstract
Research in the field of speech emotion recognition (SER) relies on the availability of comprehensive datasets to make it possible to design accurate emotion detection models. This study introduces the Multimodal Emotion Recognition and Sentiment Analysis (MERSA) dataset, which includes both natural and scripted speech recordings, transcribed text, physiological data, and self-reported emotional surveys from 150 participants collected over a two-week period. This work also presents a novel emotion recognition approach that uses a transformer-based model, integrating pre-trained wav2vec 2.0 and BERT for feature extractions and additional LSTM layers to learn hidden representations from fused representations from speech and text. Our model predicts emotions on dimensions of arousal, valence, and dominance. We trained and evaluated the model on the MSP-PODCAST dataset and achieved competitive results from the best-performing model regarding the concordance correlation coefficient (CCC). Further, this paper demonstrates the effectiveness of this model through crossdomain evaluations on both IEMOCAP and MERSA datasets.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext fdea3c50-e5e6-40b3-b82e-7c817bd7c7c7Builds on2
- wav2vec 2.0: A Framework for Self-Supervised Learning of Speech RepresentationsAlexei Baevski, Yuhao Zhou, Abdelrahman Mohamed, Michael AuliNeurIPS 2020 · 9,451 citations
- CH-SIMS: A Chinese Multimodal Sentiment Analysis Dataset with Fine-grained Annotation of ModalityWenmeng Yu, Hua Xu, Fanyang Meng, Yilin Zhu et al.ACL 2020 · 376 citations
Related papers
- Unsupervised Domain Adaptation Integrating Transformer and Mutual Information for Cross-Corpus Speech Emotion RecognitionShiqing Zhang, Ruixin Liu, Yijiao Yang, Xiaoming Zhao et al.ACM MM 2022 · 19 citations
- CM-BERT: Cross-Modal BERT for Text-Audio Sentiment AnalysisKaicheng Yang, Hua Xu, Kai GaoACM MM 2020 · 129 citations
- Unified Speech-Text Pre-training for Speech Translation and RecognitionYun Tang, Hongyu Gong, Ning Dong, Changhan Wang et al.ACL 2022 · 104 citations
- Multimodal Adaptive Emotion Transformer with Flexible Modality Inputs on A Novel Dataset with Continuous LabelsWei-Bang Jiang, Xuan-Hao Liu, Wei-Long Zheng, Bao-Liang LuACM MM 2023 · 44 citations
- WhiSPA: Semantically and Psychologically Aligned Whisper with Self-Supervised Contrastive and Student-Teacher LearningRajath Rao, Adithya V. Ganesan, Oscar N. E. Kjell, Jonah Luby et al.ACL 2025 · 4 citations
