Knowing When to Quit: Probabilistic Early Exits for Speech Separation Networks
Kenny Falkær Olsen, Mads Østergaard, Karl Ulbæk, Søren Føns Nielsen, Rasmus Malik Høegh Lindrup, Bjørn Sand Jensen, Morten Mørup
Abstract
In recent years, deep learning-based single-channel speech separation has improved considerably, in large part driven by increasingly compute- and parameter-efficient neural network architectures. Most such architectures are, however, designed with a fixed compute and parameter budget and consequently cannot scale to varying compute demands or resources, which limits their use in embedded and heterogeneous devices such as mobile phones and hearables. To enable such use-cases we design a neural network architecture for speech separation and enhancement capable of early-exit, and we propose an uncertainty-aware probabilistic framework to jointly model the clean speech signal and error variance which we use to derive probabilistic early-exit conditions in terms of desired signal-to-noise ratios. We evaluate our methods on both speech separation and enhancement tasks where we demonstrate that early-exit capabilities can be introduced without compromising reconstruction, and that when trained on variable-length audio our early-exit conditions are well-calibrated and lead to considerable compute savings when used to dynamically scale compute at test time while remaining directly interpretable.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on11
- DiffWave: A Versatile Diffusion Model for Audio SynthesisZhifeng Kong, Wei Ping, Jiaji Huang, Kexin Zhao et al.ICLR 2021 · 1,902 citations
- Going deeper with Image TransformersHugo Touvron, Matthieu Cord, Alexandre Sablayrolles, Gabriel Synnaeve et al.ICCV 2021 · 1,279 citations
- Universally Slimmable Networks and Improved Training TechniquesJiahui Yu, Thomas S. HuangICCV 2019 · 444 citations
- Stable and low-precision training for large-scale vision-language modelsMitchell Wortsman, Tim Dettmers, Luke Zettlemoyer, Ari Morcos et al.NeurIPS 2023 · 101 citations
- Hydra: Bidirectional State Space Models Through Generalized Matrix MixersSukjun Hwang, Aakash Sunil Lahoti, Ratish Puduppully, Tri Dao et al.NeurIPS 2024 · 54 citations
Related papers
- Wireless Hearables With Programmable Speech AI AcceleratorsMalek Itani, Tuochao Chen, Arun Raghavan, Gavriel Kohlberg et al.MobiCom 2025 · 4 citations
- AccCall: Enhancing Real-time Phone Call Quality with Smartphone's Built-in AccelerometerLei Wang, Xingwei Wang, Xi Zhang, Xiaolei Ma et al.UbiComp 2025 · 3 citations
- SepPrune: Structured Pruning for Efficient Deep Speech SeparationYuqi Li, Kai Li, Xin Yin, Zhifei Yang et al.AAAI 2026 · 4 citations
- mmWave-Aided Unified Speech Enhancement and Separation without Speaker Count PriorDachao Han, Teng Huang, Han Ding, Cui Zhao et al.INFOCOM 2026 · 1 citation
- Conditioning Sequence-to-sequence Networks with Learned ActivationsAlberto Gil Couto Pimentel Ramos, Abhinav Mehrotra, Nicholas Donald Lane, Sourav BhattacharyaICLR 2022 · 11 citations
