SEMOUR: A Scripted Emotional Speech Repository for Urdu
Nimra Zaheer, Obaid Ullah Ahmad, Ammar Ahmed, Muhammad Shehryar Khan, Mudassir Shabbir
摘要
Designing reliable Speech Emotion Recognition systems is a complex task that inevitably requires sufficient data for training purposes. Such extensive datasets are currently available in only a few languages, including English, German, and Italian. In this paper, we present SEMOUR, the first scripted database of emotion-tagged speech in the Urdu language, to design an Urdu Speech Recognition System. Our gender-balanced dataset contains 15, 040 unique instances recorded by eight professional actors eliciting a syntactically complex script. The dataset is phonetically balanced, and reliably exhibits a varied set of emotions as marked by the high agreement scores among human raters in experiments. We also provide various baseline speech emotion prediction scores on the database, which could be used for various applications like personalized robot assistants, diagnosis of psychological disorders, and getting feedback from a low-tech-enabled population, etc. On a random test sample, our model correctly predicts an emotion with a state-of-the-art 92% accuracy.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper1
相关 Paper
- Hate-Speech and Offensive Language Detection in Roman UrduHammad Rizwan, Muhammad Haroon Shakeel, Asim KarimEMNLP 2020 · 被引用 97 次
- The MERSA Dataset and a Transformer-Based Approach for Speech Emotion RecognitionEnshi Zhang, Rafael Trujillo, Christian PoellabauerACL 2024
- Multi-Label Compound Expression Recognition: C-EXPR Database & NetworkDimitrios KolliasCVPR 2023
- CMU-MOSEAS: A Multimodal Language Dataset for Spanish, Portuguese, German and FrenchAmirAli Bagher Zadeh, Yansheng Cao, Smon Hessner, Paul Pu Liang 等EMNLP 2020 · 被引用 43 次
- ECAvatar: 3D Avatar Facial Animation with Controllable Identity and EmotionMinjing Yu, Delong Pang, Ziwen Kang, Zhiyao Sun 等ACM MM 2024 · 被引用 4 次
