Prosody as Supervision: Bridging the Non-Verbal-Verbal for Multilingual Speech Emotion Recognition
Girish, Mohd Mujtaba Akhtar, Muskaan Singh
摘要
In this work, we introduce a paralinguistic supervision paradigm for low-resource multilingual speech emotion recognition (LRM-SER) that leverages non-verbal vocalizations to exploit prosody-centric emotion cues. Unlike conventional SER systems that rely heavily on labeled verbal speech and suffer from poor cross-lingual transfer, our approach reformulates LRM-SER as non-verbal-to-verbal transfer, where supervision from a labeled non-verbal source domain is adapted to unlabeled verbal speech across multiple target languages. To this end, we propose NOVA ARC, a geometry-aware framework that models affective structure in the Poincaré ball, discretizes paralinguistic patterns via a hyperbolic vector-quantized prosody codebook, and captures emotion intensity through a hyperbolic emotion lens. For unsupervised adaptation, NOVA-ARC performs optimal transport based prototype alignment between source emotion prototypes and target utterances, inducing soft supervision for unlabeled speech while being stabilized through consistency regularization. Experiments show that NOVA-ARC delivers the strongest performance under both non-verbal-to-verbal adaptation and the complementary verbal-to-verbal transfer setting, consistently outperforming Euclidean counterparts and strong SSL baselines. To the best of our knowledge, this work is the first to move beyond verbal-speech-centric supervision by introducing a non-verbal-to-verbal transfer paradigm for SER.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper2
- wav2vec 2.0: A Framework for Self-Supervised Learning of Speech RepresentationsAlexei Baevski, Yuhao Zhou, Abdelrahman Mohamed, Michael AuliNeurIPS 2020 · 被引用 9,451 次
- N-CORE: N-View Consistency Regularization for Disentangled Representation Learning in Nonverbal VocalizationsSiddhant Bikram Shah, Kristina T. JohnsonEMNLP 2025
相关 Paper
- On the Emotion Understanding of Synthesized SpeechYuan Ge, Haishu Zhao, Aokai Hao, Junxiang Zhang 等ACL 2026 · 被引用 1 次
- Emo-DNA: Emotion Decoupling and Alignment Learning for Cross-Corpus Speech Emotion RecognitionJiaxin Ye, Yujie Wei, Xin-Cheng Wen, Chenglong Ma 等ACM MM 2023 · 被引用 6 次
- Cross-Lingual Unsupervised Sentiment Classification with Multi-View Transfer LearningHongliang Fei, Ping LiACL 2020 · 被引用 40 次
- SAAML: A Framework for Semi-supervised Affective Adaptation via Metric LearningMinh Tran, Yelin Kim, Che-Chun Su, Cheng-Hao Kuo 等ACM MM 2023 · 被引用 5 次
- VAEmo: Efficient Representation Learning for Visual-Audio Emotion With Knowledge InjectionHao Cheng, Zhiwei Zhao, Yichao He, Zhenzhen Hu 等ACM MM 2025 · 被引用 9 次
