Prosody as Supervision: Bridging the Non-Verbal-Verbal for Multilingual Speech Emotion Recognition
Girish, Mohd Mujtaba Akhtar, Muskaan Singh
Abstract
In this work, we introduce a paralinguistic supervision paradigm for low-resource multilingual speech emotion recognition (LRM-SER) that leverages non-verbal vocalizations to exploit prosody-centric emotion cues. Unlike conventional SER systems that rely heavily on labeled verbal speech and suffer from poor cross-lingual transfer, our approach reformulates LRM-SER as non-verbal-to-verbal transfer, where supervision from a labeled non-verbal source domain is adapted to unlabeled verbal speech across multiple target languages. To this end, we propose NOVA ARC, a geometry-aware framework that models affective structure in the Poincaré ball, discretizes paralinguistic patterns via a hyperbolic vector-quantized prosody codebook, and captures emotion intensity through a hyperbolic emotion lens. For unsupervised adaptation, NOVA-ARC performs optimal transport based prototype alignment between source emotion prototypes and target utterances, inducing soft supervision for unlabeled speech while being stabilized through consistency regularization. Experiments show that NOVA-ARC delivers the strongest performance under both non-verbal-to-verbal adaptation and the complementary verbal-to-verbal transfer setting, consistently outperforming Euclidean counterparts and strong SSL baselines. To the best of our knowledge, this work is the first to move beyond verbal-speech-centric supervision by introducing a non-verbal-to-verbal transfer paradigm for SER.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 7a5b3a70-363b-4483-a47e-ca3c51cb96cbBuilds on2
- wav2vec 2.0: A Framework for Self-Supervised Learning of Speech RepresentationsAlexei Baevski, Yuhao Zhou, Abdelrahman Mohamed, Michael AuliNeurIPS 2020 · 9,451 citations
- N-CORE: N-View Consistency Regularization for Disentangled Representation Learning in Nonverbal VocalizationsSiddhant Bikram Shah, Kristina T. JohnsonEMNLP 2025
Related papers
- On the Emotion Understanding of Synthesized SpeechYuan Ge, Haishu Zhao, Aokai Hao, Junxiang Zhang et al.ACL 2026 · 1 citation
- Emo-DNA: Emotion Decoupling and Alignment Learning for Cross-Corpus Speech Emotion RecognitionJiaxin Ye, Yujie Wei, Xin-Cheng Wen, Chenglong Ma et al.ACM MM 2023 · 6 citations
- Cross-Lingual Unsupervised Sentiment Classification with Multi-View Transfer LearningHongliang Fei, Ping LiACL 2020 · 40 citations
- SAAML: A Framework for Semi-supervised Affective Adaptation via Metric LearningMinh Tran, Yelin Kim, Che-Chun Su, Cheng-Hao Kuo et al.ACM MM 2023 · 5 citations
- VAEmo: Efficient Representation Learning for Visual-Audio Emotion With Knowledge InjectionHao Cheng, Zhiwei Zhao, Yichao He, Zhenzhen Hu et al.ACM MM 2025 · 9 citations
