Conditioning Sequence-to-sequence Networks with Learned Activations
Alberto Gil Couto Pimentel Ramos, Abhinav Mehrotra, Nicholas Donald Lane, Sourav Bhattacharya
Abstract
Conditional neural networks play an important role in a number of sequence-to-sequence modeling tasks, including personalized sound enhancement (PSE), speaker dependent automatic speech recognition (ASR), and generative modeling such as text-to-speech synthesis. In conditional neural networks, the output of a model is often influenced by a conditioning vector, in addition to the input. Common approaches of conditioning include input concatenation or modulation with the conditioning vector, which comes at a cost of increased model size. In this work, we introduce a novel approach of neural network conditioning by learning intermediate layer activations based on the conditioning vector. We systematically explore and show that learned activation functions can produce conditional models with comparable or better quality, while decreasing model sizes, thus making them ideal candidates for resource-efficient on-device deployment. As exemplary target use-cases we consider (i) the task of PSE as a pre-processing technique for improving telephony or pre-trained ASR performance under noise, and (ii) personalized ASR in single speaker scenarios. We find that conditioning via activation function learning is an effective modeling strategy, suggesting a broad applicability of the proposed technique across a number of application domains.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Cited by top-tier papers1
Ask how each one uses itRelated papers
- Knowing When to Quit: Probabilistic Early Exits for Speech Separation NetworksKenny Falkær Olsen, Mads Østergaard, Karl Ulbæk, Søren Føns Nielsen et al.ICLR 2026 · 1 citation
- SeDepTTS: Enhancing the Naturalness via Semantic Dependency and Local Convolution for Text-to-Speech SynthesisChenglong Jiang, Ying Gao, Wing W. Y. Ng, Jiyong Zhou et al.AAAI 2023 · 4 citations
- AdaSpeech: Adaptive Text to Speech for Custom VoiceMingjian Chen, Xu Tan, Bohan Li, Yanqing Liu et al.ICLR 2021 · 79 citations
- Sequence to Multi-Sequence Learning via Conditional Chain Mapping for Mixture SignalsJing Shi, Xuankai Chang, Pengcheng Guo, Shinji Watanabe et al.NeurIPS 2020 · 29 citations
- Interactive Speech and Noise Modeling for Speech EnhancementChengyu Zheng, Xiulian Peng, Yuan Zhang, Sriram Srinivasan et al.AAAI 2021 · 112 citations
