Difference in Task Performance on Sparse Speech Representations
Wenjie Peng, Chen Chen, Thomas Hain
Abstract
Learning speech representations that are useful for a variety of downstream tasks has received considerable attention, due to the outstanding properties of Self-Supervised Learning (SSL) trained models. Despite advancements in modelling methods, understanding the difference in task performance on representations is limited. Mainly motivated by the nofree-lunch theorem and speech production, this work investigates changes in task performance in sparse speech representations, providing interpretability analysis under the Information Bottleneck (IB) framework. Autoencoders with varying sparsity levels were trained using three SSL features, and evaluated on six tasks of SUPERB: Speech Enhancement (SE), Speaker Identification (SID), Speech Emotion Recognition (SER), Phone Recognition (PR), Automatic Speech Recognition (ASR) and Slot Filling (SF). Experiments show that: 1) different tasks manifest different degrees of sensitivity to the sparsity levels; 2) the optimal sparsity level for task performance varies; 3) the choice of SSL features has a limited impact on most tasks but with an exception of PR; 4) overall PR and ASR require more preservation of relevant information about the labels, while SID and SER demand more compression of irrelevant information, where the input quality can shift this trade-off to some degree. These findings can contribute to the design of a universal sparse speech representation learner.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 2a19f097-f459-4a0b-b014-cd7c9f43e9deBuilds on6
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- data2vec: A General Framework for Self-supervised Learning in Speech, Vision and LanguageAlexei Baevski, Wei-Ning Hsu, Qiantong Xu, Arun Babu et al.ICML 2022 · 1,123 citations
- Sparse Autoencoders Find Highly Interpretable Features in Language ModelsRobert Huben, Hoagy Cunningham, Logan Riggs Smith, Aidan Ewart et al.ICLR 2024 · 1,072 citations
- Masked Autoencoders that ListenPo-Yao Huang, Hu Xu, Juncheng Li, Alexei Baevski et al.NeurIPS 2022 · 524 citations
- DinoSR: Self-Distillation and Online Clustering for Self-supervised Speech Representation LearningAlexander H. Liu, Heng-Jui Chang, Michael Auli, Wei-Ning Hsu et al.NeurIPS 2023 · 51 citations
Related papers
- Self-supervised Neural Factor Analysis for Disentangling Utterance-level Speech RepresentationsWeiwei Lin, Chenhang He, Man-Wai Mak, Youzhi TuICML 2023 · 6 citations
- CA-SSLR: Condition-Aware Self-Supervised Learning Representation for Generalized Speech ProcessingYen-Ju Lu, Jing Liu, Thomas Thebaud, Laureano Moro-Velázquez et al.NeurIPS 2024 · 5 citations
- Introducing Semantics into Speech EncodersDerek Xu, Shuyan Dong, Changhan Wang, Suyoun Kim et al.ACL 2023 · 2 citations
- SUPERB-SG: Enhanced Speech processing Universal PERformance Benchmark for Semantic and Generative CapabilitiesHsiang-Sheng Tsai, Heng-Jui Chang, Wen-Chin Huang, Zili Huang et al.ACL 2022 · 130 citations
- ContentVec: An Improved Self-Supervised Speech Representation by Disentangling SpeakersKaizhi Qian, Yang Zhang, Heting Gao, Junrui Ni et al.ICML 2022 · 157 citations
